> ## Documentation Index
> Fetch the complete documentation index at: https://docs.caveman.so/llms.txt
> Use this file to discover all available pages before exploring further.

# Workloads, tasks, traces, and labels in Caveman Cloud

> Learn how Caveman groups your traffic into workloads and tasks, what traces and spans capture, and how labels keep attribution clean.

Everything that flows through the Caveman gateway belongs to a workload. A workload is an observed population of traffic from one application, agent, or service. Caveman groups that traffic into tasks, records each interaction as a trace with spans, and labels everything so you can attribute cost, quality, and behavior to the right source.

## What a workload is

A workload is the unit of observation in Caveman Cloud. It represents a stream of requests that share a common shape: the same tools, prompts, or agent purpose. You can inspect workloads in the console to see cost per task, model mix, error rates, and coverage gaps.

A workload can be **registered** (declared through the SDK or console) or **observed** (discovered automatically from traffic that carries identifying labels). Registered and observed workloads keep separate identities but link together so you can trace registered intent to observed reality.

## Labels: how traffic identifies itself

Every request through the gateway should carry labels that tell Caveman which workload it belongs to. Labels are for telemetry and attribution only; they do not grant access or determine policy.

| Label | Purpose | Example |
| - | - | - |
| `x-cave-agent` | Names the agent or application sending the request | `support-agent` |
| `x-cave-workflow` | Names the workflow or task within that agent | `resolve-ticket` |

Session IDs, provided by the client or generated by the gateway, group related requests into a single conversation or task run. Use consistent labels so the console can aggregate costs and behavior accurately.

<CodeGroup>
  ```ts TypeScript theme={null}
  import { gatewayConfig } from "@caveman-ai/sdk";

  const client = new OpenAI({
    apiKey: process.env.CAVE_API_KEY,
    baseURL: `${process.env.CAVE_GATEWAY_URL}/openai/v1`,
    ...gatewayConfig({
      agent: "support-agent",
      workflow: "resolve-ticket",
    }),
  });
  ```

  ```python Python theme={null}
  from caveman_cloud import gateway_config

  config = gateway_config(
      api_key=os.environ["CAVE_API_KEY"],
      base_url=f"{os.environ['CAVE_GATEWAY_URL']}/openai/v1",
      agent="support-agent",
      workflow="resolve-ticket",
  )
  client = openai.OpenAI(api_key="unused", **config)
  ```

  ```bash curl theme={null}
  curl ${CAVE_GATEWAY_URL}/openai/v1/chat/completions \
    -H "authorization: Bearer ${CAVE_API_KEY}" \
    -H "x-cave-agent: support-agent" \
    -H "x-cave-workflow: resolve-ticket" \
    -H "x-cave-upstream-key: ${OPENAI_API_KEY}" \
    -H "content-type: application/json" \
    -d '{"model":"gpt-4o","messages":[{"role":"user","content":"hi"}]}'
  ```
</CodeGroup>

## Tasks and sessions

A **task** is one run of a workload. It includes everything that shares a session ID, or the same key and signature within the gateway's idle window. Cost per task is the objective Caveman optimizes for: it captures the total spend for a complete job, not just a single request.

Cost per task is not the same as cost per successful task. A task that retries or errors still accrues cost, and Caveman records that honestly.

## Traces and spans

Every request through the gateway generates a trace. A trace contains spans that represent individual operations:

* The gateway admission and routing span
* The provider call span
* Any optimization spans (compression, caching, routing)
* Tool execution spans, when tools are called

Spans carry timing, token counts, cost, status codes, and the basis labels that say how each number was produced. You can inspect traces in the console or query them with SQL through the `requests` and `spans` tables.

## Population: where traffic comes from

Every trace is tagged with a population class that says where the traffic originated:

| Population | Meaning |
| - | - |
| `agent` | Production agent or API traffic |
| `engineer` | Wrapped coding-agent traffic, recognized by a harness marker |
| `unknown` | No derivable identity |

Only `agent` traffic enters the continuous improvement loop. Caveman's own evaluation traffic is kept separate and is never mixed with your production workload for decision making.

## Registered versus observed workloads

When you register a workload, you declare its expected identity: agent name, tool set, and baseline model. The gateway compares observed traffic against that registration to detect drift: new tools, model switches, or cost shifts.

Observed workloads are discovered from traffic labels alone. They give you a ground-truth view of what is actually running, even if it was never registered. The console links the two so you can see whether your declared workloads match your real traffic.

## Cost per task and coverage

The console shows cost per task for each workload. This number is measured spend (list price times provider-reported tokens) divided by the number of tasks completed in the window. It is an honest average, not a projection:

* Incomplete usage rows are excluded from cost
* Unpriced models contribute zero to the numerator and are flagged
* Tasks with zero observed cost are still counted in the denominator

Coverage tells you what share of your workload traffic has complete, catalog-priced telemetry. Low coverage means some traffic is missing, unpriced, or failed, and the cost per task figure may understate reality.

## How to query workloads

You can query workload data through the Caveman SQL tables. Relevant tables include:

* `requests`: one row per gateway request, with labels, cost, tokens, and optimization outcomes
* `spans`: structured timing and attribution for each operation
* `tool_events`: tool calls and results

The `x-cave-agent` and `x-cave-workflow` columns let you group and filter by workload. See [Query with SQL](/guides/query-with-sql) for the full table reference.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.