Skip to main content
The Caveman gateway can pass requests through byte-identical or apply eligible optimizations on the fly. You control the behavior of every single call with request headers. This guide explains how to opt in, opt out, read the receipt, and understand which optimizers exist and when they run.

Record mode vs optimizations

By default, the gateway runs eligible byte-safe optimizations (S0 and S1). Behavioral and structural changes require opt-in and may need evaluation gates. A single header overrides the behavior for one request.

The x-cave-optimize header

x-cave-optimize accepts comma-separated tokens. Mixing opt-up and opt-down in one header is allowed; opt-down wins when the same optimization is named.
Unknown tokens return 400 cave_invalid_optimize_override. The gateway never guesses what you meant.

SDK per-request control

The TypeScript and Python SDKs translate options into the same header bytes.
A request never grants itself permission. The gateway checks the project policy, capability, and any required eval gate before applying an opt-in.

Response receipts and headers

Every response carries disclosure headers describing what happened. Read them to distinguish applied and denied optimizations.

Parse receipts with the SDK

A header the gateway did not send stays absent in the receipt. The receipt never fills in a default value.

Which optimizers exist

Optimizers are grouped by safety class. The gateway only runs an optimizer when its preconditions are met and the request or project policy allows it.

S0/S1 byte-safe (no model-visible change)

S1 behavioral (can change output, eval-gated)

Enabling the policy flag alone does nothing for eval-gated optimizers. The eval gate must be cleared first through an experiment that proves task success is unchanged under the change.

S2 structural (needs SDK cooperation)

  • toon-reencoding sends x-cave-optimize: compress=lossless to rewrite JSON tool output to a shorter form that round-trips to the same data.

S3 behavioral (eval-gated rollout)

  • cheaper-action-fit routes to a cheaper model that graded at least as well on replayed cases.
  • tool-output-size-profile and duplicate-step-profile identify structural patterns that require code changes in the agent.

Cache control per request

x-cave-cache controls the response cache (the gateway’s own stored answers, separate from provider prompt cache).

Opting out per request

Use opt-down tokens when one request must not touch a specific optimization:

What compression is measured at

The compression optimizer measured 33.2% fewer provider-reported input tokens on the wrap benchmark (95% CI 14.6% to 48.5%). That is a token-count reduction observed on one benchmark corpus, not a promise about your traffic and not a dollar figure. Gateway reporting stays labeled inferred until a verified method backs it.

Next steps