Capacity, Not Capability: AI’s Production Crisis Has Arrived
- By CDOTrends editors
- June 17, 2026

An AI request dies, the agent times out, and the completion never lands. The LLM then ghosts a customer mid-purchase. It’s a scenario that’s happening too often. And multiply these quiet little deaths across every API call your company fires off in a day, and you arrive at the number that should be keeping your chief data officer awake: roughly 1 in 20 AI model requests fail in production.
That’s the headline finding from Datadog’s State of AI Engineering 2026 report, drawn from anonymized telemetry across thousands of customers running large language models in the wild. 5% failure sounds tolerable until you remember traditional cloud services are measured in nines — 99.9, 99.99, 99.999. AI is operating two orders of magnitude below the bar that enterprises set for everything else they put in front of a customer.
And the culprit isn’t hallucinations or jailbreaks or bad prompts. Nearly 60% of those failures trace back to capacity limits — rate caps, throttled GPUs, and queued requests timing out. The models aren’t dumb; they’re stuck in traffic.
For CDOs, this issue reframes the entire AI conversation. Two years of board meetings have been spent on which model to pick. Datadog’s data says the question is already settled by exhaustion: 69% of companies now run three or more models in production, juggling OpenAI (still the leader at 63% share), Google Gemini, and Anthropic’s Claude, both up more than 20 percentage points year-over-year. Multi-model isn’t a strategy; it’s the floor.
What’s stacked on that floor is getting more expensive by the quarter. Agent framework adoption doubled in twelve months. Token volume per request doubled for median users (quadrupled if you are in the 90th percentile). Every workload is now sending more data, through more layers, across more vendors, with more steps that can quietly break or quietly bill.
“AI is starting to look a lot like the early days of cloud," says Yanbing Li, Datadog’s chief product officer. "The cloud made systems programmable but much more complex to manage. AI is now doing the same thing to the application layer." It means that the same finops reckoning that took a decade to tame on AWS is happening to inference, on a steeper curve, with non-deterministic costs no one can forecast.
This is squarely a data problem, which makes it squarely a CDO problem. Token flows are data flows. Agent workflows are pipelines with control logic written by a probability distribution. The hardest question of 2026 isn’t which frontier model to pick. It’s why a workflow stalled at step seven, which provider misbehaved, what context got passed where, what it just cost — and whether any of that is auditable when the regulator calls.
Vercel chief executive officer Guillermo Rauch puts the warning more sharply. “The next wave of agent failures won’t be about what agents can’t do but what teams can’t observe,” he says. “Unlike traditional software, agents have control flow driven by the LLM itself, making observability not just useful, but essential.”
The model is now writing the program. If you can’t watch the program execute, you don’t have governance. You have an invoice and hope.
Li returns to the line that should land hardest in any boardroom: "At scale, how you operate AI may matter more than the models you choose."
That’s the CDO mandate for 2026. Don’t just buy intelligence; buy the cameras. Because somewhere in your stack right now, 5% of your AI is quietly failing, the rest is quietly inflating, and no one downstream has the receipts.
Image credit: iStockphoto/elenabs