Your AI Coding Agents Need a Circuit Breaker, Not a Lower Budget Cap
Uber capped its engineers' agentic coding spend at $1,500 a month in mid-2026. Microsoft pulled most of its internal Claude Code licenses from the division behind Windows, Microsoft 365, and Teams around the same time. Neither company ran out of money because the tools stopped working. They ran out because the tools kept working, for hours longer than anyone intended, at a token price nobody had budgeted for.
That distinction matters. An AI coding agent doesn't usually fail by crashing. It fails by doing exactly what it was told, indefinitely, because nobody told it when to stop. A dollar cap set at the end of the month catches that bill a month after the fact. What actually stops a single run from eating an afternoon of compute is a pattern borrowed from site reliability engineering, not a lower number on a dashboard: the circuit breaker.
The 2026 budget reckoning
Uber gave roughly 5,000 engineers access to Claude Code in December 2025. By April 2026, adoption inside the company had climbed from 32% to 84%, and Uber had burned through its entire annual budget for AI coding tools in four months. CTO Praveen Neppali Naga told The Information the company was "back to the drawing board." President and COO Andrew Macdonald was blunter about the payoff, saying of the connection between rising AI spend and new consumer-facing features, "that link is not there yet." The response was a hard per-engineer cap: $1,500 a month on agentic coding tools like Cursor and Claude Code.
Microsoft's Experiences and Devices division, the team behind Windows, Microsoft 365, Outlook, and Teams, ran a similar rollout and hit the same wall. After six months at 84-95% adoption among thousands of engineers, per-user API costs landed between $500 and $2,000 a month. Microsoft canceled the majority of its internal Claude Code licenses effective June 30, 2026, and redirected engineers to a flatter-priced alternative.
Gartner's read on the trend, published that same June, is blunter still. The firm expects AI coding costs to exceed the average developer's own salary by 2028, driven by the shift from seat-based licensing to consumption-based token pricing. Gartner analyst Nitish Tyagi's conclusion names the actual failure point directly: "token discipline will not emerge through developer choice alone, as developers tend to optimize for speed and convenience over cost efficiency." Without something enforcing a limit, nothing does.
A monthly cap doesn't fix how agents actually fail
Both companies' first move was the obvious one: lower the number. It's also the wrong layer to fix the problem at. A monthly spending cap is a lagging indicator. It confirms, after the fact, that something ran too long or too wide. It does nothing while that something is actually running.
The uncomfortable detail in both stories is that neither required a bug. An agent working through a vague or open-ended task doesn't error out when it goes in circles. It keeps producing plausible-looking steps, each one a reasonable action on its own, for as long as its context budget and your patience hold out. Nothing in that loop looks wrong from inside a single tool call. It only looks wrong once you zoom out to the invoice, by which point the run is already over.
That's a reliability problem with a reliability answer, and it's one most teams running production services already know by a different name.
Circuit breakers, borrowed from systems you already trust
A circuit breaker, in the sense Netflix popularized with Hystrix over a decade ago, trips a service off once it starts failing past a threshold, instead of letting every downstream caller keep retrying into something already struggling. The same idea applies almost unchanged to an agent run: stop it automatically once it crosses a threshold decided on in advance, rather than relying on someone noticing mid-run, or an invoice arriving weeks later.
The difference between a circuit breaker and the dollar caps both Uber and Microsoft reached for first is where the check happens. A spending cap is enforced after usage is reported, often batched, often a day or more behind. A circuit breaker is enforced inline, before the next step runs, by the same system executing the agent, not by a finance dashboard watching from the outside.
The four boundaries worth enforcing
Most of the production patterns teams converged on through 2026 enforce the same four checks, regardless of which framework or vendor sits underneath:
- Iteration limits. A hard cap on how many steps or tool calls a single task can take before it's forced to stop and hand control back, independent of whether it's making progress.
- Budget ceilings. A dollar or token limit enforced at the infrastructure layer, checked before each step executes, not a number glanced at in a billing dashboard once a run has already finished.
- Failure thresholds. A count of consecutive errors or retries, or a confidence score dropping past a line, that trips the breaker the same way a rising error rate trips one in a traditional service.
- Scope enforcement. A permission boundary on what the agent can touch at all, so a looping run is bounded in what it can affect, not just in how long it can keep going.
None of these are exotic. They're the same four things worth wanting on any automated system making repeated calls against a metered resource. What's new is that the resource is an LLM API, the calls look like plausible code changes instead of obviously malformed requests, and the person who'd normally notice a service degrading is reading a diff instead of a dashboard.
Where the enforcement has to live
The mistake worth naming directly: a circuit breaker written as a prompt instruction, "stop after 10 steps," isn't a circuit breaker. It's a suggestion the model can drift past the same way it drifts past any other instruction once a task runs long enough. The enforcement has to sit outside the model's own reasoning, in the harness or orchestration layer actually executing each step, so a run gets killed regardless of how confident the agent sounds about continuing.
That's also why a monthly spending cap feels like a fix and isn't one. It limits the organization's total exposure eventually, but it does nothing for the single run that's three hours into a problem it can't solve. By the time that run hits the cap, the cost, in tokens spent and in the hours an engineer now has to spend reviewing whatever it produced, is already sunk.
What to actually do about this
Start smaller than a platform migration. Most agent CLIs and orchestration layers already expose a max-steps or max-turns setting; if it isn't set explicitly, there's no breaker running today. Pair that with a per-task budget ceiling, not just a per-engineer monthly one, so a single run has a hard stop before it becomes a story in a postmortem. Add a failure threshold, three consecutive tool errors or a dropping confidence score is a reasonable starting point, that forces a human check-in instead of another retry.
None of this requires the governance platforms now being sold around exactly this problem, though they exist for a reason once an org runs enough agents to need one. It requires treating an agent run the way any other automated process with a metered cost and an unbounded failure mode gets treated: give it a boundary before it starts, not a bill after it finishes.