Vision transformers versus CNNs on the factory floor

A comparison across 40,000 production-line images. The result depended less on architecture than on labeling.
We tested vision transformers against standard CNNs on production-line images, forty thousand photos across the dataset. The question everyone was asking: does the new architecture actually win on real work, or is the hype unfounded and overstated.
On the original labels, both approaches performed identically, within statistical noise. Accuracy was roughly the same. Inference speed differed but not dramatically. The model choice did not seem to matter much in the results.
We then relabeled the unclear cases, about six percent of the dataset. Ambiguous images where the label was genuinely uncertain and debatable. Once those were corrected properly, both model families improved. The improvement was larger than any architecture change we had tried before.
The real insight was straightforward but often missed: the choice of architecture mattered less than the quality and correctness of the training data. Better labels, more confident ground truth, clearer examples – those advantages outweighed any architectural innovation.
We have repeated the pattern on every vision project since. The first month is not spent experimenting with architectures. It is spent on the data: cleaning labels, removing ambiguous cases, getting ground truth right.
The lesson holds regardless of which architecture is currently trending or popular. Always start by getting the data right: clean labels, correct ground truth, unambiguous examples. Architectural choices matter, but only after the fundamentals are sound.