Skip to main content
Caveman Cloud routes model calls two ways. Requests that name the model auto (or cave-auto) let the project router pick a model from a pool you approve, per request. Requests that name a model keep it, unless you publish a route rule that moves a matching slice of traffic to a cheaper model under a measured, self-reverting rollout. This guide sets up both from Gateway → Router.

Before you start

  • Connect at least one provider under Gateway → Providers. The router only considers active, verified connections.
  • Publishing routing changes needs the admin or owner role.
  • Route your app or agent through the gateway with a project key. See Connect a workload.

Set up the router

1

Confirm a baseline

Open Gateway → Router and find Model pool. For each connected provider, choose a Baseline model. The baseline is fail-safe: it serves whenever nothing better passes the gates.If you have not set one, Caveman picks a provisional baseline per provider from your most-used model over the last 30 days. Responses carry x-cave-route-baseline-confirmed: false until you confirm or change it.
2

Approve the model pool

Check the models the router may choose. Each row shows input price per million tokens, context window, and an Evaluated badge when your evals have scored it.
  • Follow catalog lets new eligible models join the pool automatically.
  • Baseline only keeps every auto request on the baseline.
3

Tune advanced settings (optional)

Under Advanced:
  • Cost ↔ quality preference sets how far the router leans toward cheaper models inside the set that passed.
  • Routing policy chooses the router. The default, Per turn, moves small turns down and keeps hard turns on the baseline.
  • Route across providers (on by default) lets auto move a request between OpenAI and Anthropic when both are connected with a stored credential.
  • Cascade escalation sends the cheaper pick first and makes a second call only below a confidence threshold. It adds latency and real cost.
4

Publish

Click Review & publish. Every publish is a new policy version you can roll back. If the change would switch off something that is running, Caveman lists it and asks you to confirm.

Send requests to the router

Set the model to auto. Add x-cave-workflow to apply a task profile for that workflow.
For coding agents, set the model to auto wherever the agent lets you choose one. Agents that pin their own model keep it.

Price bands

Append a band to hold the router to a price range: auto:low, auto:medium, auto:high, auto:xhigh, or auto:max. An unknown band is rejected with cave_routing_invalid_cost_tier.

Read the decision

Every routed response names its decision: If the router has nothing to route from (no active provider, no baseline), the request fails with cave_routing_not_enabled and names what is missing. It is never forwarded under a model nobody chose.

Add task profiles

A task profile constrains auto for one slice of traffic. Under Task profiles, add a profile keyed by a workflow slug (such as support-triage), agent:<slug>, or default. Click Preview to see which candidates pass for a source provider and API shape, with quality, expected cost, and p95 latency, plus the reason each rejected model failed.

Move named-model traffic with route rules

Route rules change the model for requests that already name one. Click New route on the Router page and work through four steps.
1

When

Choose the workload (required). Optionally narrow it by requested model, tags, or router-classified tier.
2

Send to

Pick the target model and, optionally, a reasoning effort.
3

Compare

Backtest up to four candidates on the last seven days of matching saved traffic. A judge compares each answer to the current model’s. Each candidate gets a verdict: Holds, Worse, Can’t tell yet, or Judge unreliable. A verdict needs at least 30 paired cases and holds within a 2-point quality margin.Compare needs request bodies saved for the organization and replay allowed for the project.
4

Publish

The route goes live on 5% of matching traffic and ramps 5% → 25% → 50% → 98% while live quality holds and errors and p95 latency stay within guardrails. It pauses after 24 hours without a verdict and rolls back to 0% on a breach. 2% always stays on the original model so the saving stays measured.
Publishing is never blocked. A route without a holding backtest publishes with an Unchecked label. The first matching route wins, and route rules never apply to auto requests.

Let Caveman find cheaper routes

Automatic backtests compare cheaper models for you every hour: models suggested by open findings, and newly cheaper models on the same provider as a live route.
  • Monthly proving budget caps replay spend. Leave it empty for automatic (10% of the last 30 days’ spend, up to 250),orset250), or set 0 to $100,000.
  • Publish what holds: when the Automatic changes setting on Improvements is on, a backtest that holds goes live on a small share of traffic and ramps only while live quality holds. When off, it waits in the Inbox for you.

Measure the impact

Open Analytics → Routing for model mix and routed traffic, or chart x-cave-route-to outcomes on a custom dashboard. Routing savings are always inferred: priced against the baseline model at catalog rates and never promoted to verified. See Savings evidence.

Next steps