The problem: an AI bill you have to defend
AI agent workloads are expensive, and the spend is growing fast. When a board, CFO, or customer asks why the bill is that high, most teams can only point at dashboards. They can see costs, but they cannot explain them credibly, and they cannot prove a fix actually saved money. Vendors make that harder: they quote “up to 20x” savings from lab benchmarks, assume traffic that rarely matches yours, and invoice based on claims they never had to prove against live provider data. Caveman was built for the person who needs the number to be true.The core idea: efficiency with proof
Caveman makes every AI agent think, remember and execute on a fraction of the cost, and proves it. It sits between your agent and your provider: one base-URL swap, keep your own provider key (BYOK), and Caveman compresses repetitive context, adds cache hints, and records what happens. Every claim follows an evidence ladder: measured spend at provider-reported usage and catalog list prices; inferred estimates with a sample size, confidence, and daily range, never projected to a full month; tested changes replayed against recorded work and your eval scorers; and verified savings counted only from production traffic backed by provider data. Verified today requires one of three exact methods: provider-placed cache breakpoints on Anthropic direct; Bedrock cache points on Anthropic Claude; or a provider-counted baseline delta on counted transformed requests. No method means no verified dollar. The headline stays at $0.What makes Caveman different
Honest evidence
Verified savings start at the honest $0. A PR or passing eval is a proposal, never a saving. Overlapping detectors collapse into mutex families so headroom is never double-counted.
Eval-gated changes
Every candidate change is replayed against your traffic and graded before it can be proposed. Roll out with record, replay, shadow, canary, and automatic rollback on regression.
Byte-safe by default
Originals are recoverable byte-exact. Compression runs only when you ask for it, redacted payloads are envelope-encrypted, and any request can opt into metadata-only or zero-data-retention.
Every dollar explained
Twenty-ish detectors file each dollar under a cause, rank the moves by safety class, and put a daily dollar figure on each opportunity. No projection to month, no guesswork.
One-line integration
Swap the base URL and add one header. Works with OpenAI SDK, Anthropic SDK, Vercel AI SDK, LangChain, and LiteLLM through its OpenTelemetry exporter. Your provider key stays yours.
Governance and data control
Soft and hard caps, rate limits, guardrails, audit logs, SSO, and scoped access tokens. Deploy to our hosted cloud, your VPC via Helm and Terraform, or on-prem. ZDR stores no bodies.
Self-driving AI
The Cave Agent proposes reviewable pull requests with measured deltas. A PR is a proposal, never a saving, and never merges itself silently.
Honest evidence: measured, inferred, tested, verified
Caveman reports three kinds of numbers and keeps them separate:- Measured spend: provider-reported usage multiplied by the public catalog list price. Unknown models stay unpriced at $0.
- Inferred headroom: a modeled per-day rate from your own traffic, with a daily range, sample size, and confidence. It is never projected to a month.
- Verified savings: dollars actually not paid because a Caveman transform caused a provider-measured delta. It requires provider-complete usage, catalog pricing, and an exact causal method.
Eval-gated changes with rollback
A candidate optimization goes through a safety ladder before it touches production traffic:1
Record
The gateway records traffic byte-identical and builds a test set from your real requests.
2
Replay
The candidate is replayed against recorded work and scored with your eval suite.
3
Shadow
The change runs silently on a sample of live traffic for comparison.
4
Canary
A small share of traffic gets the change while the rest stays on the original.
5
Active
Full rollout, with automatic rollback if quality regresses.
Byte-safe by default with recoverable originals
Compression runs only when a request sendsx-cave-optimize: compress. When it runs, Caveman stores the encrypted original before any lossy replacement so it can recover byte-for-byte on demand. If parsing fails or the output is not smaller, the gateway forwards the original unchanged. Provider-native cache hints never touch model-visible text.
Raw payload storage is on by default so traffic can be replayed to prove a fix, but any request can send x-cave-retention: metadata or x-cave-retention: zdr to refuse content storage. Enterprise and customer installs never share data for training.
Every dollar explained by cause
Caveman profiles your traffic with about twenty detectors that file dollars under causes: input bloat, unused cache, duplicate tool calls, oversized tool results, routing mismatches, reliability waste, and more. Detectors are collapsed into mutex families so the same headroom is never double-counted. Each finding becomes a ranked move in the Cave Plan, with a plain-language brief and a dollar band. For SQL users, therequests table exposes cost_usd, verified_savings_usd, and routing_estimate_usd, each with a basis label that says how it was produced.
Works where your agents already run
You change one line: the base URL. Keep your provider key. Caveman supports:- Coding agents: Claude Code, Codex, Gemini CLI, OpenCode. Spend tracked per person, per agent, per merged change.
- Shipped agents: OpenAI SDK, Anthropic SDK, Vercel AI SDK, LangChain. Point the client at the gateway and add
x-cave-upstream-key. - LiteLLM: use its OpenTelemetry exporter to send traces to Caveman directly or through your OTel Collector. Provider credentials stay in LiteLLM.
Governance and data control
Caveman gives you layers of control:- Budgets: soft and hard caps per project, with Inbox items and email alerts when crossed.
- Rate limits: per key and per workflow.
- Guardrails: mask or block sensitive content in prompts and responses.
- Audit log: append-only record of changes to budgets, guardrails, keys, and roles.
- SSO: SAML and OIDC single sign-on.
- RBAC: five roles owner, admin, engineer, viewer, billing with scoped access tokens that never grant more than the calling role.
Self-driving AI that proposes, never silently auto-merges
The Cave Agent is an autonomous cost engineer. It reads telemetry, diagnoses the most expensive patterns, and proposes scoped, reviewable pull requests with their own evidence bundles and measured deltas. Your team reviews and merges. A merged PR is a proposal, not a saving, until it is active on real traffic and recorded in the ledger. See Built-in AI for how the agent loop works.What Caveman is not
Not a speed play
Not a speed play
Compression through a proxy adds latency, not removes it. We never promise faster calls. We promise cheaper, proven safe. Anyone who benchmarks us for speed should still keep us for cost.
Not a router-to-cheapest
Not a router-to-cheapest
Routing-to-cheapest is OpenRouter or Martian. Caveman may add eval-gated routing later, but it does not enter as a router. The focus is efficiency with proof, not model arbitrage.
Not observability that stops at charts
Not observability that stops at charts
Helicone and Langfuse own “see your spend.” Caveman’s console exists because seeing is a doorway: every cause links to its fix, every fix to its proof. Dashboards that dead-end at charts are the thing Caveman is not.
Not a blind compressor
Not a blind compressor
Aggressive compression degrades output badly and can even raise cost via output expansion. Caveman ships only moderate, eval-passed policies. Lossy transforms store encrypted originals for recovery.
Not byte-altering without proof
Not byte-altering without proof
Lossy compression only ships per-task after an eval clears it. Everything else stays byte-exact. The gateway never guesses what you meant.
Where Caveman pays off and where it does not
Caveman pays off when input is much larger than output and the input is repetitive or bloated. It is weaker where prompts are short, outputs are long, or quality cannot be evaluated.Benchmark and inferred envelope: approximately 20-30% potential input-cost reduction on input-heavy workloads (moderate compression benchmark about 24-28%). This is NOT an invoice or verified claim and never a guarantee of what you will see.
What it looks like for each role
Developers
One-line base-URL swap. Keep your framework, keep your provider key. Traces per request, SQL access, eval-gated improvements, and the option to run everything locally without an account.
Engineering leaders
Defensible numbers for your CFO. Governance with SSO, RBAC, scoped tokens, audit logs, and deployment to your VPC. Budget caps and guardrails that do not block engineering velocity.
Finance
Verified savings backed by provider data, signed receipts, and an evidence ladder that starts at $0. No projections, no benchmark multipliers, no “up to” claims. See it, cut it, prove it.
Next steps
Quickstart
Send your first request through the Gateway and see a live trace in minutes.
Adoption Playbook
Evaluate Caveman for your workload, build trust with stakeholders, and roll out with evidence.
Security Review
Review the security posture, data handling, tenant isolation, and BYOK architecture before you connect production traffic.