Skip to main content
Caveman Cloud gives you a complete evaluation surface to prove quality before you change production policy. You build reusable datasets, bind evaluators and targets into workbenches, run frozen comparisons, and read paired verdicts. This guide walks through creating your first eval from real traffic, running it, and using the result in CI.

Choose the right evaluation shape

Caveman Cloud supports four evaluation patterns. Pick the one that matches what you need to prove.
Optimizer experiments are a different lifecycle: they use evidence to evaluate an optimization. Creating an evaluation workbench does not publish a runtime policy.

Build a reusable dataset

Open Evals → Library in the console, or discover the datasets.create operation with the Cloud CLI. A dataset has columns for input, expected output, and metadata. Rows carry stable row_key identities. Preserve trace IDs or fixture source references so you can trace a row back to the traffic that produced it. Label synthetic cases separately from real traffic. A trace draft contains observations, not independently correct answers, so supply expectations from your task contract or review process.
1

Discover the create schema

2

Build a manifest

A manifest contains schema_version: 1, kind: "datasets", and spec matching the discovered create input.
3

Import rows

Use --rows rows.jsonl to import rows into a manifest with empty document.rows. Each JSONL line must carry its explicit row identity.
4

Save the returned manifest and revision

For updates, pull the current version, edit, validate, diff, then push again. A stale revision rejects the write.
Captured payload access remains permission and retention controlled. Missing payloads cannot be recovered from token counts or trace metadata. Keep secrets and sensitive captures out of source control.

Save evaluators and bind targets

Create reusable evaluators in Evals → Library, or use evaluators.create. Start with deterministic constraints: exact expected answers, output schema, forbidden content, ordered tools. A model judge needs an explicit rubric and independent examples of correct and incorrect behavior. Saved configuration is not qualification. Use evals.create to bind datasets, target definitions, evaluator references or inline graders, input mappings, and run settings.
1

Discover schemas before writing nested documents

2

Push the evaluator

3

Push the workbench

4

List targets for the workbench

Target types have different prerequisites:
  • HTTP target: requires a reachable authorized endpoint.
  • Recorded-output target: grades stored values and does not invoke live agent code.
  • Provider target: needs configured credentials and supported models. Missing access is a failed prerequisite.

Run, inspect, and compare

Submit a run, wait for completion, and read the result.
Use the same idempotency key only for an identical retry. Save the accepted run ID; waiting can reconnect without submitting another paid job. --timeout bounds one request and --wait-timeout bounds polling. Interrupting the waiter does not cancel remote work. Cancellation is an explicit evals.cancel operation.

Read the verdict

A judgment is one criterion answered on one unit. It carries a verdict, confidence, the grader that produced it, failing steps by reference, and its cost basis. The ultimate verdict on a change comes from the statistics kernel over paired cases: Ship, Don’t ship, or Need more data. In CI, the three verdicts map to: Pass, Regress, or Need more data.

CLI exit codes for CI

A queued or completed job alone is not a pass. For CI, preserve the command’s exit status and archive the JSON result.

Compare compatible runs

Use distinct runs over compatible frozen rows, graders, and target definitions. The comparison reports paired sample count, exclusions, sample floor, confidence or uncertainty, quality regressions, latency, and cost basis. Missing cost remains unknown; a lower token count is not verified savings. Insufficient or confounded comparisons do not establish a winner. Changing cases or grading rules requires a new baseline.

Strict suites and CI checks on pull requests

Strict suites have named cases with inline fixture objects containing candidate and reference, plus supported graders. Their baseline grades recorded outputs.
Read the returned experiment with experiments.get and experiments.results. replace_cases replaces the whole population; keep every intended case in that request. A new suite version needs a compatible baseline.
Neither fixture baselines nor ordinary workbench results grant S3 approval or certify production savings. They are quality evidence, not rollout authorization.

Qualify a judge and monitor live traffic

A saved evaluator, a passing run, and an active monitor represent different evidence and authority.

Qualify an evaluator

Open Evals → Qualification, or inspect qualifications.sets and evaluators.calibrate through operation discovery. Select an independent reference set that predates the evaluator and fits the intended population. Preserve the labels, source lineage, evaluator version, rubric, and provider configuration. Do not write holdout labels to make the current evaluator pass. Run calibration within its cost budget and inspect agreement, sample adequacy, errors, and missing evidence. Activation uses qualifications.activate and requires its own delegated scope. A pending, stale, failed, or absent qualification must not be presented as active monitoring coverage.

Monitor live traffic

Open Evals → Monitors. Configure the population, evaluator, sampling, thresholds, and delivery target with an authorized operator. Live activation depends on current qualification, project capabilities, budgets, and provider access. Agents may read monitor state and use the granted monitors.set_status pause or resume operation. A pause stops future monitoring; it does not erase past evidence. Inspect sampled population size, exclusions, errors, and verdicts alongside cost. No recent samples means unknown current coverage, not a perfect score.

Next steps