Why AI spend is hard to govern
AI workloads differ from traditional software costs in three ways that make them difficult to budget and report:- Usage-priced, not seat-priced. You pay per token, and usage can spike unpredictably when an agent retries, a workflow loops, or a model is upgraded.
- Spread across teams. Engineering, support, and operations may each run agents with different keys, models, and contract terms. Without unified telemetry, no one sees the total.
- Vendor savings claims are inflated. Many vendors report “up to” figures from benchmarks, project them to a full month, or count correlation as causation. Those numbers rarely survive audit.
How Caveman measures spend
Caveman calculates measured spend by multiplying provider-reported token counts by public catalog list prices. This happens automatically for every request that passes through the gateway.
When a model has no catalog entry, Caveman prices it at $0 and labels it
unpriced. This lowers coverage rather than guessing a price. Coverage is the share of your traffic that carries complete, catalog-priced telemetry. Low coverage does not mean Caveman is broken; it means some traffic uses unknown models or failed to report usage. The console shows coverage per workload so you can decide whether to register missing models or investigate failures.
The four evidence labels
Caveman groups every cost and savings figure into one of four labels. Each label answers a different question and belongs in a different part of a budget or board deck.Caveman never multiplies a per-day inferred rate into a monthly projection. Inferred headroom is reported as a daily rate only.
What verified requires today
Verified savings is the only per-request causal number. It requires a proven method that connects a specific Caveman transform to a provider-measured delta on that exact request. Three methods exist today:
Each method is enforced by an exact provider and optimizer tuple. A PR or passing eval is a proposal, never a saving. Verified savings stay zero when the evidence is absent, and they can be negative when cache-write premiums exceed read savings so far.
Budgets and caps with alerts
Caveman provides two types of spend controls in Governance → Budgets:- Hard caps block traffic at the gateway when settled spend plus live reservations would exceed the limit. The gateway returns HTTP 429 with
cave_budget_exceeded. - Soft spend alerts notify you when thresholds are crossed but do not block traffic. They remain advisory and fail open.
1
Open Governance → Budgets
Navigate to the Budgets page in your project.
2
Set a hard cap or soft alert
Choose the scope (project or key), the limit, and whether it blocks traffic or only alerts.
3
Choose notification recipients
Owners and admins receive Inbox items and emails automatically. Review the alert history in the Inbox.
Chargeback and showback by team, agent, and workflow
Caveman attributes every request to dimensions you can use for internal billing or cost transparency:- Team / member (requires
billing:readpermission) - Agent (via
x-cave-agentheader) - Workflow (via
x-cave-workflowheader) - API key
- Model and provider
billing:read permission. Users with the billing role can view usage and spend figures but cannot modify keys, guardrails, or team membership. Owners and admins can manage budgets and guardrails. See Team and Admin for role details.
What to ask your engineering team for
If you are setting up chargeback or showback for the first time, ask your engineering team to:- Send agent and workflow headers on every request. The
x-cave-agentandx-cave-workflowheaders label traffic automatically in traces and spend reports. - Use separate API keys per team or environment. This makes key-level budgeting and attribution straightforward.
- Grant billing:read to finance users. Add finance stakeholders with the billing role so they can view spend breakdowns without administrative access.
- Review coverage gaps. Ask which models show as
unpricedand whether they should be registered in the catalog or investigated for missing usage.
Signed receipts
Caveman supports Ed25519 receipt export for verified savings. Receipts are hash-chained, take exact 1e-10 USD units, and include the day’s distinctcave_savings_method tags. They verify included hashes and signatures, but they do not attest to omitted tails or scopes.
To request a receipt export, contact your Caveman account representative or email contact@caveman.so.
Red flags in AI cost vendor claims
When evaluating any AI efficiency vendor, including Caveman, watch for these warning signs:Fit and honest expectations
Caveman pays off when input is much larger than output and the input is repetitive or bloated. Strong fits include RAG systems, agents with large system prompts and tool schemas, tool- or log-heavy pipelines, and high-volume support bots with long preambles. Weak fits include short prompts with long outputs, already-lean prompts, and quality-critical work that cannot be evaluated. In those cases, Caveman stays byte-safe: on parse errors or unsupported inputs, the gateway forwards the original bytes unchanged. The benchmark envelope for moderate compression on input-heavy workloads is approximately 24-28% potential input-cost reduction. This is an inferred estimate, not an invoice or verified claim. Your actual verified savings depend on your traffic shape, model mix, and which verified methods apply.Data and trust
Hosted Caveman Cloud runs on GCP in europe-west4 with ClickHouse Cloud in the same region, behind Cloudflare. Request and response bodies are kept by default, envelope-encrypted with AES-256-GCM per object and wrapped by Cloud KMS, so traffic can be replayed to prove a fix. You can switch raw payload storage off to keep metadata only. Any request can sendx-cave-retention: metadata or x-cave-retention: zdr to refuse content storage.
Enterprise and customer installs never share data for training. No training runs in production today. Subprocessors are listed at caveman.so/legal/subprocessors. Legal documents, including the DPA, are drafts; request current versions from contact@caveman.so.
Caveman is not SOC 2 certified. Do not claim SOC 2, HIPAA, or any certification on Caveman’s behalf.
Next steps
Reporting Savings
A monthly reporting playbook with SQL queries, CSV exports, and board-deck phrasing.
Savings Evidence
The full technical contract for measured, inferred, and verified savings.
Team and Admin
Configure projects, invite team members, set roles, and manage billing access.
Query with SQL
Write ClickHouse SQL against requests, spend, and savings columns.