Choose the right evaluation shape
Caveman Cloud supports four evaluation patterns. Pick the one that matches what you need to prove.Optimizer experiments are a different lifecycle: they use evidence to evaluate an optimization. Creating an evaluation workbench does not publish a runtime policy.
Build a reusable dataset
Open Evals → Library in the console, or discover thedatasets.create operation with the Cloud CLI.
A dataset has columns for input, expected output, and metadata. Rows carry stable row_key identities. Preserve trace IDs or fixture source references so you can trace a row back to the traffic that produced it. Label synthetic cases separately from real traffic. A trace draft contains observations, not independently correct answers, so supply expectations from your task contract or review process.
1
Discover the create schema
2
Build a manifest
A manifest contains
schema_version: 1, kind: "datasets", and spec matching the discovered create input.3
Import rows
Use
--rows rows.jsonl to import rows into a manifest with empty document.rows. Each JSONL line must carry its explicit row identity.4
Save the returned manifest and revision
For updates, pull the current version, edit, validate, diff, then push again. A stale revision rejects the write.
Save evaluators and bind targets
Create reusable evaluators in Evals → Library, or useevaluators.create. Start with deterministic constraints: exact expected answers, output schema, forbidden content, ordered tools. A model judge needs an explicit rubric and independent examples of correct and incorrect behavior. Saved configuration is not qualification.
Use evals.create to bind datasets, target definitions, evaluator references or inline graders, input mappings, and run settings.
1
Discover schemas before writing nested documents
2
Push the evaluator
3
Push the workbench
4
List targets for the workbench
- HTTP target: requires a reachable authorized endpoint.
- Recorded-output target: grades stored values and does not invoke live agent code.
- Provider target: needs configured credentials and supported models. Missing access is a failed prerequisite.
Run, inspect, and compare
Submit a run, wait for completion, and read the result.--timeout bounds one request and --wait-timeout bounds polling. Interrupting the waiter does not cancel remote work. Cancellation is an explicit evals.cancel operation.
Read the verdict
A judgment is one criterion answered on one unit. It carries a verdict, confidence, the grader that produced it, failing steps by reference, and its cost basis. The ultimate verdict on a change comes from the statistics kernel over paired cases: Ship, Don’t ship, or Need more data. In CI, the three verdicts map to: Pass, Regress, or Need more data.CLI exit codes for CI
A queued or completed job alone is not a pass. For CI, preserve the command’s exit status and archive the JSON result.
Compare compatible runs
Strict suites and CI checks on pull requests
Strict suites have named cases with inlinefixture objects containing candidate and reference, plus supported graders. Their baseline grades recorded outputs.
experiments.get and experiments.results. replace_cases replaces the whole population; keep every intended case in that request. A new suite version needs a compatible baseline.
Neither fixture baselines nor ordinary workbench results grant S3 approval or certify production savings. They are quality evidence, not rollout authorization.
Qualify a judge and monitor live traffic
A saved evaluator, a passing run, and an active monitor represent different evidence and authority.Qualify an evaluator
Open Evals → Qualification, or inspectqualifications.sets and evaluators.calibrate through operation discovery. Select an independent reference set that predates the evaluator and fits the intended population. Preserve the labels, source lineage, evaluator version, rubric, and provider configuration. Do not write holdout labels to make the current evaluator pass.
Run calibration within its cost budget and inspect agreement, sample adequacy, errors, and missing evidence. Activation uses qualifications.activate and requires its own delegated scope. A pending, stale, failed, or absent qualification must not be presented as active monitoring coverage.
Monitor live traffic
Open Evals → Monitors. Configure the population, evaluator, sampling, thresholds, and delivery target with an authorized operator. Live activation depends on current qualification, project capabilities, budgets, and provider access. Agents may read monitor state and use the grantedmonitors.set_status pause or resume operation. A pause stops future monitoring; it does not erase past evidence. Inspect sampled population size, exclusions, errors, and verdicts alongside cost. No recent samples means unknown current coverage, not a perfect score.