FinOps · Controls
Setting Agent Spend Ceilings
A runaway agent does not overspend gradually. It overspends in the twenty minutes before anyone reads the alert.
Key takeaways
- Alerts are not ceilings. If the mechanism cannot refuse a call, it is telling you about a loss, not preventing one.
- The dangerous failure is a loop: a retry cycle can consume a month of budget in under an hour.
- A step ceiling per run is the cheapest, most effective control and almost nobody sets one.
- Decide the behaviour at the ceiling in advance — degrade, queue, or hard stop. Silent failure mid-workflow creates exception work that costs more than the spend.
- Ceilings belong at the gateway, next to policy. Anywhere else the agent can bypass them.
The short answer
Enforce four ceilings, all at the call path, none by alerting: per-run, per-agent-per-hour, per-workflow-per-day, and a per-run step limit. Monthly budget alerts are reporting, not control — they fire after the spend has happened. A ceiling has to be checked before the call is made and must degrade the workload deliberately when reached, because an agent that silently stops mid-workflow is often more expensive than one that overspends.
Why budget alerts fail
Cloud FinOps trained everyone to set budget alerts at 50%, 80% and 100%. That works when spend accrues over days from provisioned resources. Agent spend does not behave that way.
- 01The failure mode is a loop, not a trendAn agent that re-plans on a failure it cannot resolve will retry as fast as rate limits allow. That is a vertical line on the cost graph, not a slope, and threshold alerts assume a slope.
- 02Provider billing data lagsUsage often appears in billing APIs after a delay. An alert reading lagged data can arrive after the incident is over.
- 03An alert has no ability to refuseBy construction it notifies a human who then has to find and stop the thing. The gap between the two is where the money goes.
- 04Monthly windows hide daily disastersA workflow that burns its entire month in one afternoon sits comfortably inside a monthly budget until the month ends.
The distinction to hold onto
Budgets are for planning and accountability. Ceilings are for refusal. You need both, and conflating them is why teams believe they have a control when they have a report.
The four ceilings
Each catches a different failure. Together they bound both the incident and the month.
| Ceiling | Bounds | Catches |
|---|---|---|
| Steps per run | Max tool + model calls in one execution | Loops — the most common runaway |
| Cost per run | Max spend for one outcome | Expensive single runs and context blowups |
| Agent per hour | Rolling spend for one agent identity | A misconfigured or compromised agent |
| Workflow per day | Rolling spend across all agents in a workflow | Volume surges and duplicated work |
Scroll the table horizontally on narrow screens.
The step ceiling is the one to implement first. It needs no pricing data, catches the highest-frequency failure, and a sensible value is obvious from observed traffic: take the 99th percentile of steps per successful run and add headroom. Anything beyond that is not working, it is spinning.
A worked set for one workflow
workflow: invoice-to-pay steps per run max 40 (p99 observed: 23) cost per run max $0.60 (p99 observed: $0.21) agent per hour max $25 workflow per day max $600 (avg observed: $430) at ceiling: degrade → queue → alert owner
What happens at the ceiling
The part that gets skipped, and the part that determines whether the control is tolerable. An agent that stops mid-workflow with no state leaves work for a human to reconstruct — frequently costing more than the spend you saved.
Degrade first
Drop to a cheaper route, reduce retrieval breadth, disable optional enrichment. The outcome still completes, at lower cost and slightly lower quality.
Then queue
Hold the run with its state intact for the next window. Recoverable, no human reconstruction, and the work is not lost.
Then refuse cleanly
If it must stop, stop at a defined boundary with a machine-readable reason and a resumable state — not halfway through a multi-step commit.
Always notify an owner
A named human, with the run ID and what was refused. A ceiling nobody hears about will be raised by whoever hits it first.
The exception to all of it
Hard-stop immediately, with no degradation path, when the ceiling breach coincides with a policy denial spike. That combination is the signature of something abusing the agent rather than the agent doing expensive legitimate work. Cost controls and security controls should be able to see each other.
Where to enforce
- ›Inside the agent framework — bypassed by any other code path
- ›In the provider dashboard — no per-agent or per-workflow view
- ›In a nightly job — hours too late
- ›As a monthly budget alert — reporting, not refusal
- ›At the gateway, before the call is issued
- ›Next to the policy decision, sharing the same identity
- ›Reading live counters, not billed usage
- ›Returning a structured refusal the agent can handle
Enforcement needs the same three things policy needs: per-agent identity, position in the call path, and a live view of the session. That is why ceilings belong with the policy layer rather than in a finance tool — see AI gateway vs API gateway. Attribution comes from agent cost tracking; without it a ceiling cannot know whose spend it is counting.
Choosing the numbers
Derive from observation, review monthly, and expect to move them.
- Run two weeks with counters recorded and nothing enforced. You need the distribution before you can bound it.
- Set step and cost-per-run ceilings at roughly 2–3× the observed p99. Tight enough to catch a loop, loose enough not to fire on legitimate hard cases.
- Set hourly and daily ceilings from average observed spend plus a genuine growth allowance — not from the budget you were given.
- Track ceiling-hit rate per agent. A rate above about 1% of runs means the ceiling is mis-set, not that the agents are misbehaving.
- Re-derive after any material prompt, model or workflow change; all three shift the distribution.
And record every ceiling hit in the same execution record as everything else. “This run was cut short by a cost ceiling” is an important fact when somebody later asks why an invoice was not processed.
The product
Budget control as an enforcement outcome
Barzel Central Gateway treats budget control as one of six enforcement outcomes alongside denial, redaction, approval and quorum — evaluated per call with the same identity as policy. FinOps Atlas holds the attribution.
Key terms
- Spend ceiling
- A hard limit checked before a call is issued, which refuses or degrades the call when reached.
- Step ceiling
- A maximum number of model and tool calls permitted within one agent run.
- Graceful degradation
- Completing an outcome at reduced cost or quality rather than failing when a limit is reached.
- Ceiling-hit rate
- The share of runs that reach a ceiling, used to judge whether the ceiling is set correctly.
- Runaway loop
- A retry cycle that consumes budget at maximum rate without producing an outcome.
Frequently asked questions
Are budget alerts enough to control AI agent spend?
No. An alert notifies a human after spend has occurred, reads billing data that often lags, and cannot refuse a call. Agent overspend typically arrives as a loop that consumes a month of budget in under an hour.
What is the most important AI agent spend ceiling?
A step ceiling per run — a maximum number of model and tool calls in one execution. It needs no pricing data, catches runaway loops, and a sound value is derivable from the 99th percentile of steps per successful run.
What should happen when an agent hits a ceiling?
Degrade first — cheaper route, narrower retrieval — then queue the run with its state intact, then refuse cleanly at a defined boundary with a resumable state. Always notify a named owner with the run ID.
Where should ceilings be enforced?
At the gateway, alongside the policy decision, reading live counters rather than billed usage. Inside the agent framework they are bypassable; in a provider dashboard there is no per-agent or per-workflow view.
How high should a ceiling be set?
Roughly two to three times the observed 99th percentile for per-run limits, and average observed spend plus a real growth allowance for hourly and daily limits. If more than about 1% of runs hit a ceiling, the ceiling is mis-set.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.
- 01 · FinOps Foundation FinOps Framework ↗ Capability model and phases referenced throughout.
- 02 · FinOps Foundation FOCUS — FinOps Open Cost and Usage Specification ↗ Open schema for normalising cost and usage data across providers.
- 03 · Google Google SRE Book — Service Level Objectives ↗ Where the SLI/SLO/error-budget vocabulary used here comes from.
- 04 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
- 05 · NIST AI Risk Management Framework ↗ The govern / map / measure / manage structure these controls map onto.
Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.
Reference documentation
Want the specification rather than the argument?
Budget control as one of six enforcement outcomes, applied at the gateway rather than found on an invoice.
Free tier · 1,000 calls/mo · on the MCPize marketplace
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.