Skip to main content
Everything that flows through the Caveman gateway belongs to a workload. A workload is an observed population of traffic from one application, agent, or service. Caveman groups that traffic into tasks, records each interaction as a trace with spans, and labels everything so you can attribute cost, quality, and behavior to the right source.

What a workload is

A workload is the unit of observation in Caveman Cloud. It represents a stream of requests that share a common shape: the same tools, prompts, or agent purpose. You can inspect workloads in the console to see cost per task, model mix, error rates, and coverage gaps. A workload can be registered (declared through the SDK or console) or observed (discovered automatically from traffic that carries identifying labels). Registered and observed workloads keep separate identities but link together so you can trace registered intent to observed reality.

Labels: how traffic identifies itself

Every request through the gateway should carry labels that tell Caveman which workload it belongs to. Labels are for telemetry and attribution only; they do not grant access or determine policy. Session IDs, provided by the client or generated by the gateway, group related requests into a single conversation or task run. Use consistent labels so the console can aggregate costs and behavior accurately.

Tasks and sessions

A task is one run of a workload. It includes everything that shares a session ID, or the same key and signature within the gateway’s idle window. Cost per task is the objective Caveman optimizes for: it captures the total spend for a complete job, not just a single request. Cost per task is not the same as cost per successful task. A task that retries or errors still accrues cost, and Caveman records that honestly.

Traces and spans

Every request through the gateway generates a trace. A trace contains spans that represent individual operations:
  • The gateway admission and routing span
  • The provider call span
  • Any optimization spans (compression, caching, routing)
  • Tool execution spans, when tools are called
Spans carry timing, token counts, cost, status codes, and the basis labels that say how each number was produced. You can inspect traces in the console or query them with SQL through the requests and spans tables.

Population: where traffic comes from

Every trace is tagged with a population class that says where the traffic originated: Only agent traffic enters the continuous improvement loop. Caveman’s own evaluation traffic is kept separate and is never mixed with your production workload for decision making.

Registered versus observed workloads

When you register a workload, you declare its expected identity: agent name, tool set, and baseline model. The gateway compares observed traffic against that registration to detect drift: new tools, model switches, or cost shifts. Observed workloads are discovered from traffic labels alone. They give you a ground-truth view of what is actually running, even if it was never registered. The console links the two so you can see whether your declared workloads match your real traffic.

Cost per task and coverage

The console shows cost per task for each workload. This number is measured spend (list price times provider-reported tokens) divided by the number of tasks completed in the window. It is an honest average, not a projection:
  • Incomplete usage rows are excluded from cost
  • Unpriced models contribute zero to the numerator and are flagged
  • Tasks with zero observed cost are still counted in the denominator
Coverage tells you what share of your workload traffic has complete, catalog-priced telemetry. Low coverage means some traffic is missing, unpriced, or failed, and the cost per task figure may understate reality.

How to query workloads

You can query workload data through the Caveman SQL tables. Relevant tables include:
  • requests: one row per gateway request, with labels, cost, tokens, and optimization outcomes
  • spans: structured timing and attribution for each operation
  • tool_events: tool calls and results
The x-cave-agent and x-cave-workflow columns let you group and filter by workload. See Query with SQL for the full table reference.