Open weights in the enterprise: the real argument isn't price

Almost every client asking for their own model is asking for data control, not savings. It helps to say so out loud.
At mid-range volumes, serving your own model usually costs more than the API once you count GPUs, on-call and upgrades. A Lambda instance running GPT calls costs less than a reserved GPU that sits idle half the time. The economics only shift when you have predictable, sustained load.
We ran the math for a few clients with regular batch jobs. The candidates looked promising on paper. But they'd also need to maintain the infrastructure, update the model when new versions ship, and have an on-call rotation. The total cost of ownership often exceeded what they'd pay for the API.
Where there's no argument is data that cannot leave the building: healthcare, defense, certain government contracts. These aren't optimization decisions. They're compliance decisions, and they have a price. The client isn't choosing because the math is better. They're choosing because the regulation is clear.
For those cases the cost is real but inevitable. Running Llama on a private infrastructure is the only path. We help them build it, but we frame it differently: this is a regulatory requirement with infrastructure and staffing costs, not a secret path to cheaper AI.
We frame it as a compliance decision with a price tag, not an optimization. The conversations get much shorter. Leadership stops looking for the secret savings and starts budgeting for the control.