5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

FinOps · Controls

Setting Agent Spend Ceilings

A runaway agent does not overspend gradually. It overspends in the twenty minutes before anyone reads the alert.

By Mark Alex, Founder Published 19 Aug 2026 9 min read
Part of the cluster FinOps for the AI Era →

Key takeaways

  • Alerts are not ceilings. If the mechanism cannot refuse a call, it is telling you about a loss, not preventing one.
  • The dangerous failure is a loop: a retry cycle can consume a month of budget in under an hour.
  • A step ceiling per run is the cheapest, most effective control and almost nobody sets one.
  • Decide the behaviour at the ceiling in advance — degrade, queue, or hard stop. Silent failure mid-workflow creates exception work that costs more than the spend.
  • Ceilings belong at the gateway, next to policy. Anywhere else the agent can bypass them.

The short answer

Enforce four ceilings, all at the call path, none by alerting: per-run, per-agent-per-hour, per-workflow-per-day, and a per-run step limit. Monthly budget alerts are reporting, not control — they fire after the spend has happened. A ceiling has to be checked before the call is made and must degrade the workload deliberately when reached, because an agent that silently stops mid-workflow is often more expensive than one that overspends.

Why budget alerts fail

Cloud FinOps trained everyone to set budget alerts at 50%, 80% and 100%. That works when spend accrues over days from provisioned resources. Agent spend does not behave that way.

  1. 01The failure mode is a loop, not a trendAn agent that re-plans on a failure it cannot resolve will retry as fast as rate limits allow. That is a vertical line on the cost graph, not a slope, and threshold alerts assume a slope.
  2. 02Provider billing data lagsUsage often appears in billing APIs after a delay. An alert reading lagged data can arrive after the incident is over.
  3. 03An alert has no ability to refuseBy construction it notifies a human who then has to find and stop the thing. The gap between the two is where the money goes.
  4. 04Monthly windows hide daily disastersA workflow that burns its entire month in one afternoon sits comfortably inside a monthly budget until the month ends.

The distinction to hold onto

Budgets are for planning and accountability. Ceilings are for refusal. You need both, and conflating them is why teams believe they have a control when they have a report.

The four ceilings

Each catches a different failure. Together they bound both the incident and the month.

CeilingBoundsCatches
Steps per runMax tool + model calls in one executionLoops — the most common runaway
Cost per runMax spend for one outcomeExpensive single runs and context blowups
Agent per hourRolling spend for one agent identityA misconfigured or compromised agent
Workflow per dayRolling spend across all agents in a workflowVolume surges and duplicated work

Scroll the table horizontally on narrow screens.

The step ceiling is the one to implement first. It needs no pricing data, catches the highest-frequency failure, and a sensible value is obvious from observed traffic: take the 99th percentile of steps per successful run and add headroom. Anything beyond that is not working, it is spinning.

A worked set for one workflow

workflow: invoice-to-pay

steps per run          max   40   (p99 observed: 23)
cost per run           max $0.60   (p99 observed: $0.21)
agent per hour         max  $25
workflow per day       max $600   (avg observed: $430)

at ceiling: degrade → queue → alert owner

What happens at the ceiling

The part that gets skipped, and the part that determines whether the control is tolerable. An agent that stops mid-workflow with no state leaves work for a human to reconstruct — frequently costing more than the spend you saved.

Stage 01

Degrade first

Drop to a cheaper route, reduce retrieval breadth, disable optional enrichment. The outcome still completes, at lower cost and slightly lower quality.

Stage 02

Then queue

Hold the run with its state intact for the next window. Recoverable, no human reconstruction, and the work is not lost.

Stage 03

Then refuse cleanly

If it must stop, stop at a defined boundary with a machine-readable reason and a resumable state — not halfway through a multi-step commit.

Stage 04

Always notify an owner

A named human, with the run ID and what was refused. A ceiling nobody hears about will be raised by whoever hits it first.

The exception to all of it

Hard-stop immediately, with no degradation path, when the ceiling breach coincides with a policy denial spike. That combination is the signature of something abusing the agent rather than the agent doing expensive legitimate work. Cost controls and security controls should be able to see each other.

Where to enforce

Will not hold
  • ›Inside the agent framework — bypassed by any other code path
  • ›In the provider dashboard — no per-agent or per-workflow view
  • ›In a nightly job — hours too late
  • ›As a monthly budget alert — reporting, not refusal
Holds
  • ›At the gateway, before the call is issued
  • ›Next to the policy decision, sharing the same identity
  • ›Reading live counters, not billed usage
  • ›Returning a structured refusal the agent can handle

Enforcement needs the same three things policy needs: per-agent identity, position in the call path, and a live view of the session. That is why ceilings belong with the policy layer rather than in a finance tool — see AI gateway vs API gateway. Attribution comes from agent cost tracking; without it a ceiling cannot know whose spend it is counting.

Choosing the numbers

Derive from observation, review monthly, and expect to move them.

  • Run two weeks with counters recorded and nothing enforced. You need the distribution before you can bound it.
  • Set step and cost-per-run ceilings at roughly 2–3× the observed p99. Tight enough to catch a loop, loose enough not to fire on legitimate hard cases.
  • Set hourly and daily ceilings from average observed spend plus a genuine growth allowance — not from the budget you were given.
  • Track ceiling-hit rate per agent. A rate above about 1% of runs means the ceiling is mis-set, not that the agents are misbehaving.
  • Re-derive after any material prompt, model or workflow change; all three shift the distribution.

And record every ceiling hit in the same execution record as everything else. “This run was cut short by a cost ceiling” is an important fact when somebody later asks why an invoice was not processed.

The product

Budget control as an enforcement outcome

Barzel Central Gateway treats budget control as one of six enforcement outcomes alongside denial, redaction, approval and quorum — evaluated per call with the same identity as policy. FinOps Atlas holds the attribution.

Key terms

Spend ceiling
A hard limit checked before a call is issued, which refuses or degrades the call when reached.
Step ceiling
A maximum number of model and tool calls permitted within one agent run.
Graceful degradation
Completing an outcome at reduced cost or quality rather than failing when a limit is reached.
Ceiling-hit rate
The share of runs that reach a ceiling, used to judge whether the ceiling is set correctly.
Runaway loop
A retry cycle that consumes budget at maximum rate without producing an outcome.

Frequently asked questions

Are budget alerts enough to control AI agent spend?

No. An alert notifies a human after spend has occurred, reads billing data that often lags, and cannot refuse a call. Agent overspend typically arrives as a loop that consumes a month of budget in under an hour.

What is the most important AI agent spend ceiling?

A step ceiling per run — a maximum number of model and tool calls in one execution. It needs no pricing data, catches runaway loops, and a sound value is derivable from the 99th percentile of steps per successful run.

What should happen when an agent hits a ceiling?

Degrade first — cheaper route, narrower retrieval — then queue the run with its state intact, then refuse cleanly at a defined boundary with a resumable state. Always notify a named owner with the run ID.

Where should ceilings be enforced?

At the gateway, alongside the policy decision, reading live counters rather than billed usage. Inside the agent framework they are bypassable; in a provider dashboard there is no per-agent or per-workflow view.

How high should a ceiling be set?

Roughly two to three times the observed 99th percentile for per-run limits, and average observed spend plus a real growth allowance for hourly and daily limits. If more than about 1% of runs hit a ceiling, the ceiling is mis-set.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.

  1. 01 · FinOps Foundation FinOps Framework ↗ Capability model and phases referenced throughout.
  2. 02 · FinOps Foundation FOCUS — FinOps Open Cost and Usage Specification ↗ Open schema for normalising cost and usage data across providers.
  3. 03 · Google Google SRE Book — Service Level Objectives ↗ Where the SLI/SLO/error-budget vocabulary used here comes from.
  4. 04 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
  5. 05 · NIST AI Risk Management Framework ↗ The govern / map / measure / manage structure these controls map onto.

Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.

Reference documentation

Want the specification rather than the argument?

Budget control as one of six enforcement outcomes, applied at the gateway rather than found on an invoice.

Free tier · 1,000 calls/mo · on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.