GPT-4: what actually changes for an enterprise project

Fewer silly errors and longer context. That reopens use cases we shelved six months ago, and leaves others closed.
The big jump is following complex instructions without drifting halfway through a document. When you ask GPT-3.5 to parse a specification with multiple constraints and extract consistent fields, it loses the thread partway in. GPT-4 maintains context and completes the entire task coherently. The difference compounds over longer documents and complicated logic.
That stability opened doors to projects that looked risky on an older model. We could ask it to transform data according to a set of rules, apply those rules consistently across a corpus, and trust the output without hand-checking every case. The reliability changed the economics.
We revisited three pilots that had stalled on quality. Two cleared the client's acceptance bar; the third still needed data the company had never structured. The model improved, but the data didn't. Without clean input, no model would have solved it.
In one case the company had never built a taxonomy for their documents. They had a folder and a hope. The model could help, but first someone had to define what was actually in there. That work was outside the model's scope and it was real.
A better model doesn't fix an information problem. We've repeated that sentence all year. Every proposal now starts by auditing whether the company owns the data it's asking the model to use. If they don't, the model upgrade is theatre.