Governance surfaces
Open Governance in the console for these controls:
Budget and guardrail decisions require
policy:draft scope. Testing guardrails requires policy:draft as well. Audit log reads need auditlog:read.
Budget caps per project
Hard budgets
A hard budget blocks traffic at admission when the cap is reached. The gateway reserves a conservative maximum before each upstream call using counted input tokens orlen(body)/2 input tokens, the request’s own output cap or 8192 output tokens, at public catalog list prices. If the projected spend would exceed the cap, the gateway returns 429 Too Many Requests with error code cave_budget_exceeded.
Clients that honor
x-should-retry (such as OpenCode) stop instead of retrying, because an exhausted cap is not transient.
Soft budgets
A soft budget does not block requests. When the threshold is crossed, the gateway serves the request and adds alert headers so your monitoring or automation can act:
Soft threshold crossings also feed the budgets dashboard and email alerts. You configure both values under Governance → Budgets.
Budget bypass
Keys with theadmin:bypass_budget scope skip hard and soft budget checks. Use this only for operational or emergency access, and audit its use regularly.
Rate limits
Per-key rate limits
You can set a requests-per-minute (RPM) limit and a concurrency limit on individual API keys under Governance → Keys. When a key exceeds its RPM, the gateway returns429 with error code cave_key_rate_limit_exceeded. The response includes Retry-After so callers can back off. No x-should-retry header is set because rate limits are explicitly transient.
A key’s rate limit narrows what the project already allows; it never widens the project limit. If the project RPM is 1 and the key RPM is 100, the project limit still governs.
Per-workflow spend rate limits
Spend rate limits protect against runaway costs on a rolling window. You can configure:- Per-key USD threshold: total spend allowed for one API key in the window
- Per-workflow USD threshold: total spend allowed for one workflow in the window
- Window seconds: the rolling window duration
- Soft block: when true, exceeded scopes return an advisory refusal; when false, only headers are added
429 with error code cave_spend_rate_soft_limit_exceeded. The response includes:
Spend rate accrual is idempotent: duplicate request IDs do not double-count. If Valkey is unavailable, spend rate checks fail open so traffic is not blocked by an infrastructure issue.
Guardrails
Guardrails screen request and response content for patterns you define. A rule has a kind, a mode (when it runs), an action (what happens on a match), and optional match parameters.Rule kinds
Rule modes
- Pre-call: screens the request before it reaches the provider
- Post-call: screens the provider’s response before it reaches the client
Rule actions
- Block: refuses the request or response with
400 Bad Request - Mask: rewrites the matched text in place (for example,
alice@example.combecomes[REDACTED:email]) and forwards the sanitized body
record) body is not allowed because the gateway promised to forward the caller’s exact bytes. If a mask rule fires while the project is in record mode, the gateway blocks instead.
Block behavior
When a guardrail blocks, the response is:- Status
400 Bad Request - Error code
cave_guardrail_blockedfor pre-call blocks - Error code
cave_guardrail_blocked_responsefor post-call blocks - Details naming the
guardrail,kind, andmode - The refusal message does not echo the matched text or credential
Post-call stream screening
For streaming responses, post-call guardrails buffer and screen each chunk. If a secret or blocked pattern is split across fragments, the joined text is still matched. If the stream is too large to buffer, the gateway fails closed and refuses the response. Unreadable stream formats (for example, Bedrock eventstream frames) also fail closed.Testing guardrails before publishing
Test a guardrail configuration against sample text before you publish it to the gateway. Use the console or the API:policy:draft scope. It returns which rules would match and what action each would take, without publishing any configuration.
Publish guardrails with a PUT to /projects/{projectId}/guardrails. Re-read the delivered state after publishing: a saved configuration is not proof the gateway has applied it.
Audit log verification
Every change to budgets, guardrails, keys, roles, and provider connections is recorded in the append-only audit log under Governance → Audit. Each row contains:- Actor (who made the change)
- Resource (what was changed)
- Time and outcome
- Source IP and session context where available
auditlog:read, which is owner/admin scoped.
Recommended setups
Per environment
Create separate projects or API keys for development, staging, and production. Apply strict hard budgets and guardrails in production. Use permissive soft budgets in development so experiments are not interrupted, but still receive alerts when spend is unexpectedly high.Per team
Issue scoped access tokens that narrow an integration to a subset of the caller’s role. Authorization is the intersection of role and scope, so a scope never grants authority the role lacks. For example, give a CI service a token withevals:run and sql:read but not policy:draft or billing:write.
Per workflow
Label traffic withx-cave-workflow so spend rate limits and tracing are meaningful. A workflow label lets you set a per-workflow spend rate threshold and query costs by workflow in SQL.
RBAC quick reference
Caveman Cloud uses five roles. Sensitive actions are narrowed to owner and admin:
Publishing aggressive (S2/S3) policies and approving S3 experiments is owner/admin only. Raw payload read is owner/admin only. Connecting a repository (which grants write agency for auto-PR agents) is owner/admin only.
Next steps
Team and Admin
Invite members, assign roles, and configure SSO.
Control Optimizations
Set per-request optimization headers and read receipt headers.
Query with SQL
Query guardrail hits, budget status, and spend by workflow in ClickHouse.
Troubleshooting
Fix common budget, authentication, and trace issues.