Skip to main content
Caveman Cloud budgets stop runaway model spend at the gateway, before a request reaches the provider. You can set a monthly cap for a whole project, a cap per API key, and short-window spend-rate quotas that catch bursts within minutes. Configure project budgets under Governance → Budgets and key budgets under Governance → Keys.

Budget types at a glance

All amounts are in US dollars at public catalog list prices. Set any value to 0 to turn it off.

Set a monthly project budget

1

Open Governance → Budgets

Select the project. The Monthly budget panel shows gateway spend so far this UTC month, the current caps, and the amount reserved in flight.
2

Set a soft cap

Enter the amount at which you want a warning. Crossing it notifies owners and admins through the Inbox and email, and fires any budget.alert webhook, once per period. Traffic keeps flowing.
3

Set a hard cap

Enter the amount at which traffic must stop. Once committed spend plus in-flight reservations reaches the cap, new requests are refused until the 1st of next month (UTC) or until you raise the cap.
4

Save limits

Click Save limits. The console confirms when the policy has been delivered to the gateway. If it says delivery is pending, the change is saved but not yet enforced.
Every soft or hard cap crossing appears in the Cap crossings panel and in the Inbox.

How hard caps avoid overshoot

Before forwarding a request, the gateway reserves its maximum possible cost: the counted input tokens plus the largest allowed output, at catalog price. After the provider answers, it settles the real amount and releases the rest. Because reservations count against the cap, many concurrent requests cannot collectively blow past it. Requests to an unpriced model reserve a fixed ceiling.
Setting a sensible max_tokens on your requests makes reservations tighter, so you can run closer to a hard cap without early refusals.

Spend-rate quotas

Monthly caps protect the month; spend-rate quotas protect the next few minutes. They watch rolling spend so a looping agent is caught long before it dents the monthly budget. Recent threshold crossings stay in the Spend-rate alerts panel for 7 days.
Spend-rate quotas are advisory. The estimate is a rolling lower bound, in-flight provider work can still spend, and the quota fails open if its counter store is unavailable. Use a hard cap when you need a guaranteed stop.

Handle budget responses in your app

Hard cap reached

The gateway answers 429 with code cave_budget_exceeded and the header x-should-retry: false. Do not retry; the request will keep failing until the window resets or the cap is raised.
Response body

Soft cap crossed

The request succeeds and the response carries warning headers you can log or surface:

Spend-rate quota exceeded

When refusal is on, the gateway answers 429 with code cave_spend_rate_soft_limit_exceeded. Back off and retry later, or route the work to a different key.

Spend accounting

The Spend accounting panel puts three figures side by side:
  • Measured layer spend: provider-reported usage priced at catalog list rates. This is what caps count.
  • Imported spend: spend reported by external systems. Shown for correlation only and excluded from caps.
  • Verified savings: provider-measured savings that a Caveman change caused. See Savings evidence.

Who can change budgets

Owners, Admins, and Engineers can edit budgets. Keys with the admin:bypass_budget scope are never metered or blocked, so keep them rare. Budget changes are recorded in Governance → Audit.