Skip to main content
Adopting Caveman Cloud is a four-week exercise with clear owners, milestones, and exit criteria each week. This playbook is designed for three audiences working together: developers who integrate and operate the gateway, platform and engineering leaders who govern the rollout, and finance stakeholders who need defensible numbers. Each week builds on the last: you connect traffic, diagnose what it costs, prove that changes are safe, and report verified savings that can stand up to scrutiny.
This plan assumes you already have a Caveman Cloud account. If not, start with the Quickstart to sign in, create a project, and send your first request.

RACI at a glance

R = Responsible, A = Accountable, C = Consulted, I = Informed

Week 1: Connect and baseline

Your goal this week is to get one production workload routing through the Caveman gateway, with labeled traffic, budgets in place, and confirmed coverage. Nothing is optimized yet. You are in record mode, collecting an honest baseline.

Developer tasks

1

Swap the base URL

Change one production workload to route through Caveman. This is a single line change: point your OpenAI, Anthropic, or compatible client at the gateway URL with your Cave API key.
See Connect a Workload for Anthropic, Google, and curl examples.
2

Label every request

Set x-cave-agent and x-cave-workflow on every request. Labels are case-sensitive in traces and determine how traffic groups into workloads. Inconsistent casing fragments your data and makes cost attribution harder.
3

Connect coding agents (optional)

If your team uses Claude Code, Codex, Gemini CLI, or OpenCode, wrap them with caveman so their traffic is captured and attributed per person, per agent, per merged change. See Connect Coding Agent.
4

Confirm traces appear

Open Traces in the console and verify your traffic is visible. Each trace should show model, tokens, cost, latency, and the agent and workflow labels you attached.

Platform/CTO tasks

1

Set budgets

Open Governance → Budgets and configure soft spend alerts and hard caps per project. Soft alerts send an Inbox item and email when crossed. Hard caps block traffic at admission when the reservation would exceed the limit.
2

Configure guardrails

Set up guardrails that mask or block sensitive content in prompts and responses. Test them with policy.test_guardrails before publishing an authorized policy change.
3

Review access controls

Confirm team roles in Team and SSO. Five roles exist: owner, admin, engineer, viewer, billing. Raw payload read is owner/admin only. An admin cannot mint or modify owner access.

Finance tasks

1

Confirm coverage

Open Spend and check the coverage share. Coverage is the percentage of traffic with complete, catalog-priced telemetry. Low coverage means some models are unpriced or requests are missing token counts, which makes downstream savings claims unreliable.
2

Document the baseline

Record the first measured spend number. This is your baseline (list price times provider-reported tokens). It is not your provider invoice, but it is the honest starting point for every comparison.

Week 1 exit criteria

  • At least one workload routing through the gateway with consistent labels
  • Traces visible in the console with agent and workflow attribution
  • Soft budget alerts and hard caps configured
  • Coverage above 80% for the connected workload, or a plan to register missing models

Common pitfalls in Week 1


Week 2: Diagnose

Your goal this week is to understand where your money is going, read the ranked improvement opportunities, and select one to two concrete moves to prove. You will also define eval criteria so you can judge whether those moves are safe.

Developer tasks

1

Read the Cave Plan

Open Improvements → Cave Plan to see ranked moves per workload. Each move shows inferred headroom in dollars per day, broken down by safety class (S0 to S3). Start with S0/S1 moves: they are byte-safe and need no code changes.
2

Query top cost drivers with SQL

Run a query to find your most expensive workflows and models:
See Query with SQL for more examples.
3

Check fit

Caveman pays off when input is much larger than output and the input is repetitive or bloated: RAG chunks, tool schemas, system prompts, and long histories. If your workload is short prompts with long outputs, or already lean, the inferred headroom will be honest about the smaller prize.

Platform/CTO tasks

1

Pick 1-2 moves

From the Cave Plan, pick one S0/S1 move and optionally one S2 move. Good first candidates:
  • Anthropic cache breakpoints (anthropic-cache-breakpoints): byte-safe, on by default, path to verified dollars
  • Bedrock cache points (bedrock-cache-points): opt-in, also path to verified dollars
  • TOON reencoding (toon-reencoding): S2, needs x-cave-optimize: compress=lossless, eval-gated
Avoid starting with S3 behavioral changes (model routing, output caps). They require eval gates and canary rollouts before they earn anything.
2

Define eval criteria

For any move that changes model-visible bytes, define what “safe” means. Build a dataset in Evals → Library from saved traffic, or create a strict suite with inline fixtures. See Build and run evals.

Finance tasks

1

Validate inferred numbers

The Cave Plan shows daily rates, never monthly projections. Ask: does the inferred headroom align with the share of spend that is input-heavy? If most spend is on output tokens, the opportunity is smaller.

Week 2 exit criteria

  • Cave Plan reviewed and top cost drivers identified
  • 1-2 moves selected with clear owners
  • Eval criteria defined for any move above S0
  • Baseline measured spend documented for comparison

Common pitfalls in Week 2


Week 3: Prove

Your goal this week is to run experiments, clear eval gates, and earn your first verified savings. Byte-safe cache optimizations are the fastest path to verified dollars because Caveman can prove causality with provider-reported cache usage.

Developer tasks

1

Run a replay experiment

For S2/S3 moves, replay the candidate change against recorded traffic. Start the experiment from the improvement in the console, or run cvm tools list to find the experiment operations available to your role. The experiment replays saved requests with the new policy and grades outputs against your eval criteria.
2

Shadow mode

If available for your move, run in shadow mode first. The gateway executes the optimization but does not serve the result. This proves the change is safe before any user sees it.
3

Canary with eval gates

For moves that passed replay, enable them on a small share of traffic with x-cave-optimize headers or project policy. Monitor eval results continuously. If quality degrades, rollback is immediate.
4

Enable verified cache optimizations

For Anthropic-direct or Bedrock Claude traffic, ensure cache breakpoints or cache points are active. These are the only optimizations that can earn verified savings today because Caveman can prove the provider billed the cache reads or writes.

Platform/CTO tasks

1

Approve S1/S2 experiments

Publishing aggressive (S2/S3) policies and approving S3 experiments requires owner or admin role. Review the Evidence report: one claim, approach, physics proof, judged proof, evidence links, proving cost, and recommended action.
2

Monitor the ledger

Open Improvements → Verified savings daily. Verified savings start at the honest $0. They grow only when qualifying evidence exists. Cache-heavy days may show negative verified savings at first (write premiums), then turn positive as reads accumulate.

Finance tasks

1

Count verified dollars

Verified savings is the only number that can be defended in a finance readout. Inferred headroom is an estimate; verified savings is provider-grounded causal proof. If verified is still zero, document why (record mode, unpriced models, or non-qualifying traffic).

Week 3 exit criteria

  • At least one replay or shadow experiment completed
  • Eval gate cleared for any S1+ move
  • First verified savings rows visible in the ledger, or documented reason why not
  • Canary running on a bounded traffic share with rollback plan

Common pitfalls in Week 3


Week 4: Report and expand

Your goal this week is to present the first verified numbers, expand to additional workloads, and enable Automation carefully. This is where the loop becomes continuous.

Developer tasks

1

Expand to more workloads

Apply the same connect-and-label pattern to the next 1-2 workloads. Use the proven eval criteria from Week 3 where relevant.
2

Enable Automation (carefully)

From the installation settings, enable repository scans and production investigations first. Hold off on repair pull requests until your team is comfortable reviewing Evidence reports. See Set up Automation.
3

Connect a decision model (optional)

For consistent Compare and route judgment, connect a dedicated decision model in Settings → Judges. This improves the independence of eval verdicts.

Platform/CTO tasks

1

Present verified results

Summarize the first verified savings number with coverage percentage and the method that produced it (provider_causal_cache, provider_causal_cache_bedrock, or provider_counted_baseline_delta). Include the honest $0 baseline if nothing is verified yet.
2

Set a proving budget

Configure the monthly proving budget in Settings. The default is 10% of your last 30 days’ model spend, capped at 250 USD. Set it to 0 USD to turn proving off entirely.
3

Review governance changes

Audit logs record changes to budgets, guardrails, keys, and roles. Review Governance → Audit after any administrative change.

Finance tasks

1

Draft the first readout

Include three numbers with their rung labels:
  • Measured spend: baseline cost at catalog list price
  • Inferred headroom: daily opportunity rate, not a promise
  • Verified savings: causal dollars saved, with qualifying method and coverage
Do not sum inferred and verified. They answer different questions.
2

Plan the next reporting cycle

Verified savings grow as cache reads accumulate and more workloads qualify. Set a monthly cadence to review the ledger and expand the coverage percentage.

Week 4 exit criteria

  • First verified savings reported with coverage and method
  • 2-3 workloads connected and labeled
  • Automation enabled for at least repository scans
  • Finance readout delivered with honest labels

Readiness checklist

Before you start Week 1, confirm these prerequisites:
1

Account and project

You have a Caveman Cloud account, a project, and at least one Cave API key. You know your gateway URL.
2

Provider keys

You have active provider keys (OpenAI, Anthropic, Bedrock, etc.) and permission to route traffic through them.
3

Team access

Team members are invited with the right roles. Owners and admins can configure governance. Engineers can integrate and query.
4

Workload identified

You have named the first workload to connect: an agent or service with measurable traffic and clear input/output characteristics.
5

Retention policy understood

You know your organization’s payload storage setting and retention windows. Payload storage is on by default; ZDR is available per request.

Common adoption pitfalls


What comes after Week 4

The 30-day playbook gets you to your first verified number and a repeatable loop. After that:
  • Monthly: review the Cave Plan, run new experiments, and expand verified methods to more workloads
  • Quarterly: report verified savings to stakeholders with coverage trends
  • Continuously: let Automation scan repositories and traffic, but keep human approval at every gate
For deeper guidance on each surface, see:

Connect a Workload

Route production traffic with a base-URL swap and two headers.

Traces and Spend

Inspect cost, latency, and verified savings in the console.

Improvements

Review the Cave Plan, run attempts, and read Evidence reports.

Evals

Build datasets, run workbenches, and clear eval gates.

Automation

Enable continuous scans and investigations.

Savings Evidence

Understand measured, inferred, and verified labels.

Control Optimizations

Choose record mode or active optimizations per request.

Team and Admin

Configure roles, SSO, budgets, and governance.