Whisper: transcription that's finally useful

Tested on 200 hours of calls with noise, accents and crosstalk. It's the first time the output doesn't need rewriting end to end.
Whisper's real improvement was not in average-case accuracy metrics. It was in the worst cases – the overlapping speech, the accents, the background noise – which is exactly where human correction effort went previously.
We tested it on two hundred hours of real support calls. The audio was messy and realistic. People interrupted each other. Accents varied. Previous systems had lost confidence in these situations. Whisper stayed coherent throughout.
The model has clear limits that matter. Proper nouns and product codes stump it. Technical or specialized terminology outside its training set gets mangled or invented. Rather than admitting uncertainty, it confidently produces plausible sounds that fit the context.
We built a post-processing step that ran transcripts against a client dictionary: product names, account numbers, regulatory terms. Anything outside the dictionary that looked like proper noun got flagged for human review.
Most transcripts needed no human review afterward. When review was necessary, corrections went quickly because Whisper's output was close enough to guide the editing process. The gaps were small and the suggestions credible.
With transcription that is actually reliable, downstream systems become feasible. Conversation analysis, sentiment tracking, compliance auditing. Those stop being research projects and become operational tools running on real data.