What Caveman stores
The gateway records two categories of data from every request:- Metadata: request timing, token counts, cost, model name, status code, labels (
x-cave-agent,x-cave-workflow), session IDs, optimization outcomes, and trace identifiers - Payloads: request and response bodies, tool results, compression originals, and SDK artifacts
Payload storage and consent
Three consents govern what happens to your data. They are stated separately and can be managed independently:
Payload storage is the floor. Without it, neither training use nor replay consent can reach anything, because bytes that were never kept cannot be reused. Training use and replay are independent above that floor: you can allow replay for evaluation without consenting to product-wide training, or vice versa.
Replay consent is on by default for every project. A person in your organization can revoke it per project through the console or API. Free organizations may have restrictions on switching payload storage or training use off, depending on plan terms.
Retention and deletion
You control how long data is kept through retention windows you set on your organization. Retention is a policy choice, not a product limit, and the rule is the same on every plan.
Derived working data (semantic cache samples, traffic vectors, quality monitor checkpoints, router outcomes) keeps fixed system lifetimes. Shortening your request-history window also trims derived data that depends on it.
When you set a window:
- The daily retention job deletes expired rows and objects
- Deletion is permanent: lifting the window later does not restore what was deleted
- The console asks for confirmation before shortening a window
Request capture details
Request captures share chunks with earlier captures from the same project, retention class, and creation hour. A capture whose earlier shared chunks have expired can become unreadable up to an hour before its own window ends. This is a storage optimization: sharing never keeps bytes past the retention window. Capture is best-effort under load. When the gateway’s capture queue or byte budget is full, the request is still served and recorded, but its payloads are skipped and markedskipped_backpressure.