What a workload is
A workload is the unit of observation in Caveman Cloud. It represents a stream of requests that share a common shape: the same tools, prompts, or agent purpose. You can inspect workloads in the console to see cost per task, model mix, error rates, and coverage gaps. A workload can be registered (declared through the SDK or console) or observed (discovered automatically from traffic that carries identifying labels). Registered and observed workloads keep separate identities but link together so you can trace registered intent to observed reality.Labels: how traffic identifies itself
Every request through the gateway should carry labels that tell Caveman which workload it belongs to. Labels are for telemetry and attribution only; they do not grant access or determine policy.
Session IDs, provided by the client or generated by the gateway, group related requests into a single conversation or task run. Use consistent labels so the console can aggregate costs and behavior accurately.
Tasks and sessions
A task is one run of a workload. It includes everything that shares a session ID, or the same key and signature within the gateway’s idle window. Cost per task is the objective Caveman optimizes for: it captures the total spend for a complete job, not just a single request. Cost per task is not the same as cost per successful task. A task that retries or errors still accrues cost, and Caveman records that honestly.Traces and spans
Every request through the gateway generates a trace. A trace contains spans that represent individual operations:- The gateway admission and routing span
- The provider call span
- Any optimization spans (compression, caching, routing)
- Tool execution spans, when tools are called
requests and spans tables.
Population: where traffic comes from
Every trace is tagged with a population class that says where the traffic originated:
Only
agent traffic enters the continuous improvement loop. Caveman’s own evaluation traffic is kept separate and is never mixed with your production workload for decision making.
Registered versus observed workloads
When you register a workload, you declare its expected identity: agent name, tool set, and baseline model. The gateway compares observed traffic against that registration to detect drift: new tools, model switches, or cost shifts. Observed workloads are discovered from traffic labels alone. They give you a ground-truth view of what is actually running, even if it was never registered. The console links the two so you can see whether your declared workloads match your real traffic.Cost per task and coverage
The console shows cost per task for each workload. This number is measured spend (list price times provider-reported tokens) divided by the number of tasks completed in the window. It is an honest average, not a projection:- Incomplete usage rows are excluded from cost
- Unpriced models contribute zero to the numerator and are flagged
- Tasks with zero observed cost are still counted in the denominator
How to query workloads
You can query workload data through the Caveman SQL tables. Relevant tables include:requests: one row per gateway request, with labels, cost, tokens, and optimization outcomesspans: structured timing and attribution for each operationtool_events: tool calls and results
x-cave-agent and x-cave-workflow columns let you group and filter by workload. See Query with SQL for the full table reference.