1
Day 1: send one traced request
Create a Cave API key in the console, set two environment variables, and change your client’s base URL. Keep your upstream provider key in Open Traces in the console, filter by agent and workflow, and confirm the request, cost, latency, and tokens are visible.
x-cave-upstream-key or store it once in Gateway, Connections and drop the header.2
Week 1: label, control, and debug
Label traffic with The gateway answers with disclosure headers:Debug with Traces. Every request becomes a trace. Click a trace to inspect spans, see provider-reported tokens, inferred compression estimates, and the exact headers that were sent. If a trace is missing, check the auth header, gateway URL spelling, and provider path.Set retention per request with
x-cave-agent and x-cave-workflow. Labels do not grant access, but they power every report: spend per agent, per workflow, per merged PR. Choose values that match your application structure and keep casing consistent.Control optimizations per request with x-cave-optimize. Pass off to go byte-identical, compress to request compression, or no-cache to disable one optimizer while leaving others running.Use the SDK’s
parseReceipt to read them. The SDKs also translate options into the same header bytes.TypeScript
Python
x-cave-retention: metadata or x-cave-retention: zdr to refuse payload storage for a single call. This is honored on every plan.3
Month 1: query, evaluate, and automate
Query with Results include columns, rows, elapsed time, and a query digest you can cite as evidence. Agent connections and Connect your coding agent over MCP. Claude Code, Codex, or Cursor can search traces, run SQL, and read reports directly from your editor. Open Connect your coding agent in the console, approve the project binding, and verify with
cvm sql. The CLI runs read-only ClickHouse SELECT over your project’s tables.sql:read keys are limited to 101 rows, 128 KiB, and 15 seconds. People get up to 10,000 rows, 8 MiB, and 30 seconds.Run evals in CI. Author strict suites of test cases, run them with the CLI, and gate merges on the exit code. Evals clear the evidence gates that let behavioral optimizations run in production.caveman_context. The agent gets metadata reads by default; content and authoring are opt-in.Authentication
Caveman Cloud uses two separate credentials. The Cave API key authenticates you to Caveman Cloud. The upstream provider key authenticates to the model provider. The gateway readsx-cave-api-key first and falls back to authorization: Bearer <key>. For full details and common errors, see Authentication.
Per-request optimization control
Optimizations are a project-level setting, but a single request can ask for or disable specific ones. The gateway checks the project policy, capability, and evidence gate before applying anything. A request never grants itself permission.
Unknown tokens return 400
cave_invalid_optimize_override. The gateway never guesses what you meant.
SDK per-request control
Response receipts
Every response carries disclosure headers. Parse them with the SDK, or read them directly:tokensBefore and tokensAfter are inferred estimates from the compressor, not provider counts. If recoveryHandle is present, you can retrieve the byte-exact original.
Retention per request
Any request can tighten retention for that call only:
These are honored on every plan and override the project’s default for that request only. For org-wide retention settings, see Data and Privacy.
Debugging with Traces
The Traces page in the console shows every request, with model, tokens, latency, cost, and the agent and workflow labels you attached. Click a trace to inspect:- Request and response spans
- Provider-reported usage vs inferred estimates
- Applied and denied optimizations
- Compression ratios and recovery handles
- The exact headers sent
- Confirm
authorization: Bearer <Cave key>is present - Check that
CAVE_GATEWAY_URLhas no trailing slash - Verify the provider path matches your SDK (
/openai/v1,/anthropic, etc.) - Ensure
x-cave-upstream-keyis present or a stored connection is configured
Querying with cvm sql
The cvm CLI gives you read-only SQL over your project’s telemetry.
sql:read keys are limited to 101 rows, 128 KiB, and 15 seconds. People get up to 10,000 rows, 8 MiB, and 30 seconds. Results include query_digest, which you can cite as evidence.
For the full SQL guide, see Query with SQL.
Evals in CI
Evaluations are strict suites of test cases that prove a proposed change does not break quality. Run them in CI and gate merges on the result.MCP from your coding agent
Connect Claude Code, Codex, or Cursor to Caveman Cloud over MCP so the agent can inspect traces, run SQL, and read reports.- Open Connect your coding agent in the console
- Select your client and approve the project binding
- Call
caveman_contextfrom the agent to verify the connection
caveman_context, caveman_search, caveman_describe, caveman_read, caveman_write, and caveman_draft. Metadata reads are default; content reads require explicit permission. See MCP for the full tool reference and protocol details.
Coding agent routing with the open-source caveman CLI
For spend tracking per person, per agent, and per merged change, route your local coding agent through the Caveman gateway using the open-source caveman CLI. No account is needed for local tools.
aider, claude, codex, gemini, hermes, openclaw, and opencode. Each sets x-cave-agent so telemetry knows which agent made the request. For the full setup, see Connect Coding Agent.
SDK features
The TypeScript (@caveman-ai/sdk) and Python (caveman-sdk, import caveman_cloud) packages wrap the gateway’s /sdk/v1/* contract with zero runtime dependencies.
Core capabilities
- Trace workflows: wrap a unit of work with
cave.trace()and get helpers for tools, model calls, artifacts, and checkpoints - Per-request control: set
optimize,cache, andstyleoptions per call - Model routing: send
model: "cave-auto"and let the project’s router pick - Receipt parsing:
parseReceipt/parse_receiptreads disclosure headers - Compression:
cave.compress()delegates to the gateway; on failure it passes the original through unchanged - Tool search:
cave.tools()returns a catalog handle withsearch()that queries the gateway - OTel exporter:
cave.exporter()builds OTLP/JSON by hand and POSTs to/v1/traceswithout external dependencies - Shared context: hand context between agents in the same project with
cave.sharedContext.put/get - Prompt snippets:
cave.prompts.internalBrevity()builds deterministic brevity instructions
TypeScript quick example
Python quick example
cvm CLI is the terminal interface for Caveman Cloud. Install it with npm, sign in once, and run traces, SQL, evals, and Cloud operations from your shell.
For the full CLI reference, see cvm CLI.
Next steps
Quickstart
Get from zero to a live trace in minutes.
Connect a workload
Route production traffic through the gateway end to end.
Control optimizations
Choose record mode or active optimizations per request.
Evals
Build test cases and evaluate candidate changes with evidence.
Query with SQL
Run read-only SELECT over requests, spans, and tool events.
MCP
Connect your coding agent to Caveman Cloud over MCP.