RACI at a glance
Week 1: Connect and baseline
Your goal this week is to get one production workload routing through the Caveman gateway, with labeled traffic, budgets in place, and confirmed coverage. Nothing is optimized yet. You are in record mode, collecting an honest baseline.Developer tasks
Swap the base URL
Label every request
x-cave-agent and x-cave-workflow on every request. Labels are case-sensitive in traces and determine how traffic groups into workloads. Inconsistent casing fragments your data and makes cost attribution harder.Connect coding agents (optional)
caveman so their traffic is captured and attributed per person, per agent, per merged change. See Connect Coding Agent.Confirm traces appear
Platform/CTO tasks
Set budgets
Configure guardrails
policy.test_guardrails before publishing an authorized policy change.Review access controls
Finance tasks
Confirm coverage
Document the baseline
Week 1 exit criteria
- At least one workload routing through the gateway with consistent labels
- Traces visible in the console with agent and workflow attribution
- Soft budget alerts and hard caps configured
- Coverage above 80% for the connected workload, or a plan to register missing models
Common pitfalls in Week 1
Week 2: Diagnose
Your goal this week is to understand where your money is going, read the ranked improvement opportunities, and select one to two concrete moves to prove. You will also define eval criteria so you can judge whether those moves are safe.Developer tasks
Read the Cave Plan
Query top cost drivers with SQL
Check fit
Platform/CTO tasks
Pick 1-2 moves
- Anthropic cache breakpoints (
anthropic-cache-breakpoints): byte-safe, on by default, path to verified dollars - Bedrock cache points (
bedrock-cache-points): opt-in, also path to verified dollars - TOON reencoding (
toon-reencoding): S2, needsx-cave-optimize: compress=lossless, eval-gated
Define eval criteria
Finance tasks
Validate inferred numbers
Week 2 exit criteria
- Cave Plan reviewed and top cost drivers identified
- 1-2 moves selected with clear owners
- Eval criteria defined for any move above S0
- Baseline measured spend documented for comparison
Common pitfalls in Week 2
Week 3: Prove
Your goal this week is to run experiments, clear eval gates, and earn your first verified savings. Byte-safe cache optimizations are the fastest path to verified dollars because Caveman can prove causality with provider-reported cache usage.Developer tasks
Run a replay experiment
cvm tools list to find the experiment operations available to your role. The experiment replays saved requests with the new policy and grades outputs against your eval criteria.Shadow mode
Canary with eval gates
x-cave-optimize headers or project policy. Monitor eval results continuously. If quality degrades, rollback is immediate.Enable verified cache optimizations
verified savings today because Caveman can prove the provider billed the cache reads or writes.Platform/CTO tasks
Approve S1/S2 experiments
Monitor the ledger
Finance tasks
Count verified dollars
Week 3 exit criteria
- At least one replay or shadow experiment completed
- Eval gate cleared for any S1+ move
- First verified savings rows visible in the ledger, or documented reason why not
- Canary running on a bounded traffic share with rollback plan
Common pitfalls in Week 3
Week 4: Report and expand
Your goal this week is to present the first verified numbers, expand to additional workloads, and enable Automation carefully. This is where the loop becomes continuous.Developer tasks
Expand to more workloads
Enable Automation (carefully)
Connect a decision model (optional)
Platform/CTO tasks
Present verified results
provider_causal_cache, provider_causal_cache_bedrock, or provider_counted_baseline_delta). Include the honest $0 baseline if nothing is verified yet.Set a proving budget
Review governance changes
Finance tasks
Draft the first readout
- Measured spend: baseline cost at catalog list price
- Inferred headroom: daily opportunity rate, not a promise
- Verified savings: causal dollars saved, with qualifying method and coverage
Plan the next reporting cycle
Week 4 exit criteria
- First verified savings reported with coverage and method
- 2-3 workloads connected and labeled
- Automation enabled for at least repository scans
- Finance readout delivered with honest labels
Readiness checklist
Before you start Week 1, confirm these prerequisites:Account and project
Provider keys
Team access
Workload identified
Retention policy understood
Common adoption pitfalls
What comes after Week 4
The 30-day playbook gets you to your first verified number and a repeatable loop. After that:- Monthly: review the Cave Plan, run new experiments, and expand verified methods to more workloads
- Quarterly: report verified savings to stakeholders with coverage trends
- Continuously: let Automation scan repositories and traffic, but keep human approval at every gate