The usual approach is adding metadata and validation later in the data pipeline. Rather than cleaning data and losing the original context, it’s often more effective to keep as much information about the original state of the data, says David Aronchick, open-source platform Kubeflow founder, and CEO of distributed data pipeline vendor Expanso. “You can’t pursue exactly purely clean data; that’s just not possible,” he says. “As you pull data into your ML model, every line should have some mechanism saying where it came from. But that information may well be relevant down the line when you want to use that data more broadly.