The safety ladder: S0 through S4
The safety ladder is the contract between Caveman and your application. Lower tiers are safer and more automatic. Higher tiers unlock larger savings but require cooperation, evaluation, or explicit opt-in.
No optimizer above S0 runs automatically. Every change that can alter model-visible text requires either a request header or an approved evaluation gate.
Gateway optimizers: S0 and S1 byte-safe
These optimizers run inside the gateway on the upstream request only. They never change model-visible content. They add provider-native caching hints, and the response stream to your client is untouched. Each is policy-gated, idempotent, and passes the body through byte-identically on any parse problem.anthropic-cache-breakpoints
anthropic-cache-breakpoints
Provider: Anthropic (direct) Default: On for every request Opt-out: Send
x-cave-optimize: no-cacheInjects one ephemeral cache_control marker on the largest stable prefix: tools if present, otherwise system. When the same prefix repeats on later requests, Anthropic bills it at the cache_read rate instead of full input price. The marker is Caveman’s own breakpoint, so rows where it triggers cache reads qualify for verified savings under the provider_causal_cache method.openai-prompt-cache-key
openai-prompt-cache-key
Provider: OpenAI Default: On for every request Opt-out: Send
x-cave-optimize: no-cacheInjects a canonical prompt_cache_key derived from sha256(model + tools + system/instructions + session or user id) to raise prompt-cache routing affinity while keeping each routing group under OpenAI’s overflow threshold (about 15 requests per minute). OpenAI caching is automatic at 1,024-token prefixes or longer; this optimizer only improves hit rate. It never enables impossible caching, and it costs nothing if prefixes vary.Because OpenAI caches automatically, this optimizer does not place a Caveman-specific marker. Its effect is observed from provider usage (
prompt_tokens_details.cached_tokens), but it remains inferred. There is no verified method for OpenAI cache affinity today.bedrock-cache-points
bedrock-cache-points
Provider: Bedrock (Anthropic Claude) Default: Off (opt-in) Opt-in: Project policy enables it; request needs no extra headerAdds one Bedrock-native cache marker to the largest stable tools or system prefix. Caller-managed markers, models without catalog prompt-cache capability, non-pay-as-you-go auth, and malformed bodies pass through byte-identically. When enabled and eligible, qualifying cache rows use the
provider_causal_cache_bedrock verified method.Retired Gemini scaffold
gemini-explicit-cache is retained only as a historical join key and unconditional pass-through. It has no current policy, practice, experiment, proposal, or verified-savings path. Gemini’s current Interactions API supports implicit caching only. The legacy generateContent API still supports explicit cachedContents resources, but they require lifecycle management (creation, credential project, model and Vertex location, TTL, deletion, creation charge, and token-hour storage) that the gateway does not route or preserve. Requests that already reference cachedContent remain provider pass-through and are deliberately unpriced.
Behavioral optimizers: S1, eval-gated
Unlike the byte-safe optimizers above, these can change model output. They run only when both a policy flag is enabled and an eval gate is cleared. The eval gate is set to true only after an experiment proves task success is unchanged under the proposed change.
Enabling the policy flag alone does nothing. The eval gate protects the hard-task tail where dialing reasoning down or capping output can silently degrade quality. In dry-run observations, verbose workloads show recorded
output_tokens dropping (for example, 800 to 512) while task output is preserved. A request whose eval gate is not cleared is passed through untouched. This is quality evidence, not a verified billing claim.
Compression: S4 with a lossless tier
Compression is the most powerful optimizer Caveman offers, and it is gated accordingly. The full S4 engine elides content and stores the original under a content-addressed handle so it can be retrieved byte-for-byte later. Because it changes what the model sees, it runs only when the request asks for it.How to request compression
Send the header on a single request:Lossless tier: when recovery cannot run
The S4 engine requires a CCR retrieve loop and stored originals. Where that loop cannot run, acompress ask falls back to a lossless tier instead of being denied. This applies to streaming, the OpenAI Responses API, Gemini, and metadata-only or ZDR retention traffic.
The lossless tier rewrites tool output only:
- An Anthropic
tool_result, OpenAItoolmessage, Responsesfunction_call_output, or GeminifunctionResponsethat contains a single JSON object or array becomes TOON when TOON decodes back to exactly the same value. - Otherwise it becomes compact JSON.
- Whichever is shorter in bytes wins. Non-JSON text, and JSON that neither form can shorten, is sent byte-identically.
caveman_retrieve tool is added, and no prefix-cache write occurs.
Request lossless explicitly
x-cave-optimize: compress=lossless admits only the lossless tier on every route where compress works, streaming or not. It never runs the S4 engine or adds the retrieve tool. When both compress and compress=lossless appear on one header, the lossless ask wins. The disclosure header reads x-cave-optimize-applied: compress=lossless.
Verified compression savings
Compression savings become verified through theprovider_counted_baseline_delta method. The transform runs only when the request asks, so the ask is the grant. Where Caveman’s cache breakpoint owns a row (because it placed the marker), the cache delta owns the row and the counted delta is dropped, an under-claim. Where the caller manages caching (for example, Claude Code), the breakpoint does not run, so the counted delta mints there. Counted rows that read the cache are priced at the cache-read rate when the delta is a saving, and at the input rate when it is a loss.
What the benchmark says
The compression optimizer measured 33.2% fewer provider-reported input tokens on the wrap benchmark (95% confidence interval 14.6% to 48.5%). That is a token-count reduction on one benchmark corpus, not a promise about your traffic, and not a dollar figure. Gateway reporting stays labeledinferred until a verified method backs it.
Routing: cave-auto
The cave-auto model name asks the project router to pick one eligible model. As of the per-request optimization decision, the model name decides routing: cave-auto (and auto, auto:<tier>, cost_tier) is always answered by the router. Nothing switches it off (no capability, opt-in, optimizer flag, eval gate, canary state, or no-route token). The only refusals are structural: no active provider connection, no baseline model, an endpoint without a rewritable model, a lossless or locally transformed request, or an unknown tier.
A named model is never re-chosen by the router. Route rules a person added for a named model still apply unless the request says no-route.
Default projects with no configured baseline now receive a provisional baseline picked from the project’s most-used named model over the last 30 days, or the provider’s reviewed default. The response carries x-cave-route-baseline and x-cave-route-baseline-confirmed: false until someone confirms or changes it on the Router page.
Routing savings are always inferred. They are priced against the baseline model and never promoted to verified. See the full header reference on Control Optimizations, and set up the router in Model Routing.
Detectors and mutex families
The Caveman worker profiles traffic and writesopportunities rows through detectors. Every detector dollar figure is inferred: it is a catalog-priced counterfactual, not a provider invoice. To prevent double-counting, overlapping detectors are collapsed into mutex families so the same headroom is never counted twice.
The four families
Input bloat
Input bloat
Historical mutex identity retains retired money IDs. Current telemetry emits zero-dollar
context-window-profile, tool-catalog-profile, and exploration-load-profile outside the mutex. Two live members remain:toon-reencoding(S2): payload-level measurement on sampled captured bodies that proves JSON tool results round-trip through TOON and are smaller.tool-output-size-profile(S3): the part of each tool result past a 2,000-estimated-token cap and its carries through the session.
Cache
Cache
provider-cache-unused(S2) andsemantic-cache-opportunity(S1) carry catalog-priced inferred headroom.prompt-prefix-stabilityandcache-miss-root-cause(F1) are report-only with zero dollars.cache-write-read-churnis a retired historical mutex identity, not a current pricing signal.
Reliability
Reliability
Old money IDs are retired. Live members:
usage-accounting-missing: live with a registered zero band.duplicate-step-profile(S3): the recorded spend of requests that only repeated a tool call within 60 seconds.
Routing
Routing
Old heuristic IDs are retired. One live member:
cheaper-action-fit(S3): sums savings only over replayed cases that graded at least as well as the recorded model. Without replay, it reprices the scope-day’s usage at the nearest cheaper catalog model (inferred, quality not yet tested).
(agent, workflow, model, day): the one with the highest base estimate. See Savings Evidence for how inferred numbers relate to verified ones.
Detector safety classes: what zero change earns
Each detector carries the safety class of the fix it recommends. The class is the honest answer to “what does capturing this cost me?” Caveman rolls the headroom up by class so an operator can see what zero app change earns versus what cooperation or a behavioral change unlocks.
Byte-safe (S0/S1) optimization captures provider-native caching. The larger structural levers (S2) need SDK cooperation. The biggest lever (S3 model/reasoning routing) can change output and sits behind the eval-gated rollout by design. The split is the truthful shape of where the money is and what each tier costs to claim.
Verified savings eligibility by optimizer
Not every optimizer can produce verified savings. Verified methods require a proven causal contract and provider-complete usage.
Verified savings can be negative. A cache write that nobody reads books a loss. A compressed request that the provider counted as larger than the original also books a loss. Days can legitimately show negative verified savings, and Caveman never floors the number to zero.
Which optimization fits my workload
Use this checklist to match your traffic shape to the optimizers that can help.Long system prompts repeated every turn
Best fit:
anthropic-cache-breakpoints or openai-prompt-cache-key Action: Nothing to change; these run by default. Opt out with no-cache if your prefixes never repeat.Large tool schemas or JSON results
Best fit:
compress=lossless or full compress Action: Send x-cave-optimize: compress on tool-heavy requests. The lossless tier handles streaming or Responses API traffic automatically.Verbose generations with runaway length
Best fit:
output-brevity Action: Enable the policy flag, then run an eval suite to clear the gate. The cap applies only when you do not set your own.Reasoning-capable models on simple tasks
Best fit:
reasoning-effort Action: Enable the policy flag, run evals, and the gate lowers thinking tokens when you do not specify an effort level.High-volume repetitive questions
Best fit: Semantic response cache Action: Send
x-cave-cache: semantic and ensure your project precision score is healthy. Shadow mode measures matches before serving.Cost concentration in one model or task
Best fit:
cheaper-action-fit and detector-driven moves Action: Inspect the Cave Plan for ranked moves per workload. Prove cheaper models with replayed evals before any rollout.Per-role guidance
- Developers
- CTOs and Engineering Leaders
- CFOs and Finance
You control every optimization per request with headers. No dashboard switch overrides your code.
- Use
x-cave-optimize: offto pass one request through unchanged. - Use
x-cave-optimize: compressto request compression on tool-heavy turns. - Use
x-cave-optimize: no-cacheto disable cache hints for one request. - Use
x-cave-optimize: compress=losslessto force the JSON-only tier. - Read
x-cave-optimize-appliedandx-cave-optimize-deniedto audit what ran.
Response receipts: how to read what ran
Every response carries disclosure headers. Parse them to distinguish applied and denied optimizations.
In the TypeScript SDK, use
parseReceipt(response.headers). In Python, use parse_receipt(response.headers). Headers the gateway did not send stay absent; the receipt never fills in a default value.