Skip to main content
Caveman Cloud is built for the person who must defend the AI budget to a CFO, a board, or a customer. It gives you spend per developer, team, agent, and workflow; ties coding sessions to merged pull requests so each PR shows its cost; and counts a saving only after production traffic proves it. This page covers the problem, the controls, the governance model, and the evaluation checklist for a pilot.

The problem: an ungoverned AI spend line item

AI spend is now a real budget line, but few teams can answer the basic questions:
  • How much did we spend last month, and on what?
  • Which team, agent, or workflow drove the increase?
  • Did that compression or routing change actually save money, or did it just look good in a dashboard?
  • If the CEO asks for proof, what can we show?
Caveman Cloud answers every question with a labeled evidence rung: measured (what the provider reported), inferred (a modeled daily rate), tested (a replayed candidate), or verified (a provider-measured delta). No number is presented without its source.

What you get org-wide

Spend per developer, team, agent, and workflow

The gateway records every request with its agent and workflow labels, so you slice spend by any dimension. Coding sessions (Claude Code, Codex, Gemini CLI, OpenCode) link to their branch and to the change that merged, so each pull request shows the sessions behind it and what they cost. The console surfaces this under Traces, Workloads, and Developers.

Cave Plan: ranked moves with daily dollar figures

Caveman reads your traffic for repeated work, bloated context, and the wrong model for the job. Each finding becomes a brief: what was found, the proposed change, and how to verify it. The brief carries an inferred daily dollar figure, a sample size, and a confidence. You choose what to act on. See Improvements for how to review proposals and Built-in AI for how the system generates them.

Evidence ladder: four labeled rungs

Every claim is labeled so you know what you are looking at: Verified savings today require one of three provider-causal methods: Anthropic-direct cache breakpoints Caveman placed, Bedrock Anthropic Claude cache points, or provider-counted baseline delta on counted transformed requests. A passing eval or a merged PR is a proposal, never a verified saving. See Savings Evidence for the full contract.

Risk model: byte-safe defaults and staged rollout

Caveman does not guess about quality. The default is byte-safe pass-through, and every optimization climbs a safety class before it can run on production traffic.

Byte-safe default

The gateway forwards your request byte-identical unless you ask for an optimization. Compression runs only when the request carries x-cave-optimize: compress. On any parse problem, unsupported input, or output that is not smaller, the gateway passes the original unchanged. Originals are recoverable byte-exact under a content-addressed handle. Record mode skips every transform.

Safety classes and eval gates

Optimizers are gated by safety class:

Staged rollout with auto-rollback

An optimizer rolls out through record, replay, shadow, canary, and active stages. Each stage has a gate, and live quality monitors sample production traffic. If a monitor detects regression, the system auto-rolls back to the prior state. The rollback is automatic; the re-activation is human. See Rollout Safety for the stage definitions.

Reviewable PRs, never silent auto-merges

When the Cave Agent proposes a code change, it holds the change and its evidence in the Inbox. A person with repository connect permission chooses Publish, and only then does Caveman’s GitHub App push the branch and open a draft PR. An agent session cannot choose Publish. Your team reviews, merges, and releases. We recommend enabling GitHub Workflow Execution Protections and leaving Caveman’s GitHub App off the allowed list, so a person on your team starts CI.

Who can approve what

Five roles control what each person can do: Sensitive scopes are narrowed: raw payload read is owner/admin only; publishing aggressive (S2/S3) policies and approving S3 experiments is owner/admin; connecting a repo is owner/admin. An admin cannot mint or modify owner access.

Governance

SSO

Sign in via email and password, Google, GitHub, or OIDC/SAML single sign-on. SAML signature verification runs in the identity service with no external auth vendor. See Team and Admin for setup steps.

RBAC and scoped tokens

Five roles define the authorization boundary. Scoped access tokens narrow an integration to a subset of the caller’s role permissions. Authorization is the intersection of role and scope: a scope never grants authority the role lacks. The scope catalogue is published in the OpenAPI specification.

Budgets and rate limits

Set soft and hard caps per project. A hard cap blocks traffic; a soft cap sends an Inbox item and an email. Rate limits apply per key and per workflow. See Budgets and Guardrails for configuration.

Guardrails

Guardrails mask or block sensitive content in prompts and responses. Project-level rules layer over the built-in floor. Test a guardrail before publishing a policy change. See Control Optimizations for how guardrails interact with optimization headers.

Audit log

An append-only audit log records changes to budgets, guardrails, keys, roles, and policies. Inspect actor, resource, time, and outcome for every change. The log is tamper-evident within the tenant boundary.

Vendor neutrality and BYOK

Caveman Cloud is bring-your-own-key. You configure your own provider credentials; we transmit requests to those providers with your credentials, as your instruction. Your relationship with the model provider is your own.

No lock-in

  • Swap the base URL back to the provider to remove Caveman from the path.
  • Local tools (caveman <agent>) need no account and keep working if Cloud is unreachable.
  • Coding-agent traffic goes straight to your provider with your key. Caveman gets a routing ask, not the request.
  • Enterprise runs in your VPC, on your KMS and storage, with content sharing locked off.

Scoped access tokens

Scoped tokens let you give an integration only the permissions it needs, narrowed from the caller’s role. This is useful for CI pipelines, third-party agents, or departmental projects.

Build vs buy

If you were to build what Caveman provides, you would need at least:
  • A provider-complete pricing catalog that stays current as models and prices change.
  • A causal savings accounting system that starts at $0 and requires provider-measured proof before minting a dollar.
  • A replay pipeline that runs recorded traffic against candidate changes with your scorers.
  • An eval gating system that blocks transforms until quality is proven unchanged.
  • Row-level tenant isolation, envelope encryption, and a retention worker with audit logging.
  • A staged rollout controller with shadow, canary, and auto-rollback.
Caveman ships these as one Helm chart with a signed release. Hosted Cloud runs them for you; customer-owned installs run them in your VPC.

Evaluation checklist for a pilot

Use this checklist to evaluate Caveman Cloud against your requirements.

KPIs to track

Run the pilot for at least two billing cycles and track these numbers:
  • Spend per merged change: total model spend divided by merged PRs in the window. This tells you if coding agents are getting more or less expensive per unit of shipped work.
  • Cost per task: for agentic workflows, divide spend by completed task count.
  • Verified savings rate: verified savings divided by measured spend on eligible traffic. A negative number is valid and honest.
  • Coverage: the share of traffic that carries complete, catalog-priced telemetry. Low coverage means some models are unpriced or some requests failed.
  • Inferred headroom per day: the modeled daily opportunity from Cave Plan. Compare it to the cost of the proving budget and the engineering time to review proposals.
  • Proposals reviewed vs merged: the ratio of improvement attempts that reached a PR to those that shipped. This measures your team’s velocity, not Caveman’s.

Next steps

Deployment Options

Hosted Cloud, your VPC on AWS or GCP, or on-prem.

Security Review

Architecture, encryption, tenant isolation, and compliance posture.

Savings Evidence

How measured, inferred, and verified numbers are produced and labeled.

Team and Admin

Projects, roles, SSO, keys, budgets, and guardrails.

Improvements

Prepare workloads, review Evidence reports, and merge proposed changes.

Budgets and Guardrails

Set soft and hard caps, rate limits, and content guardrails.