Skip to main content
Caveman Cloud is a managed LLM gateway plus workload telemetry, evaluation, and improvement tooling built for teams running AI agents and LLM traffic at scale. It connects request-level usage to workload-level decisions so you can understand costs, prove improvements, and deploy changes with evidence. Caveman Cloud is self-driving software for your LLM stack. Built-in agents observe your traffic, evaluate candidate changes, prove them with evidence, and propose improvements. You stay in control and approve what ships. The loop runs continuously: Observe, Diagnose, Generate, Evaluate, Deploy, Learn. What sets Caveman apart is that it proves its own claims. Every number carries a label (measured, inferred, tested, or verified), changes ship only after they clear eval gates, the default path is byte-safe, and verified savings start at an honest $0. Read Why Caveman for the full picture.

Choose your path

Developers

Integrate in one line, label traffic, control optimizations per request, and debug with traces, SQL, and evals.

CTOs and Platform Leads

Govern AI spend across teams, agents, and coding tools with eval-gated rollouts, RBAC, and audit.

CFOs and Finance

Understand what each number means, which savings are verified, and how to report them with confidence.

30-Day Adoption Playbook

Connect, diagnose, prove, and report in four weeks.

Security Review

Hosting, encryption, key custody, tenant isolation, and data policy.

Deployment Options

Hosted Cloud, your own VPC, on-prem, or local-only tools.

Let AI drive

Connect Your Coding Agent

Give Claude Code, Codex, or Cursor access to your project over MCP.

Automated AI Setup

Paste one prompt and let your coding agent wire up the gateway, evals, and baseline.

Built-in AI

See how Ask Caveman, Improvements, judges, and rollouts work together.

Improvements

Review agent-generated changes backed by Evidence reports before they ship.

Get started

Quickstart

Send your first request through the Gateway and see it in Traces within minutes.

Connect a Workload

Route agent or application traffic with a base-URL swap and two headers.

How It Works

Understand the Observe to Learn loop and how Caveman Cloud handles your traffic.

OpenAI Integration

Point the OpenAI SDK (TypeScript or Python) at the managed gateway.

TypeScript SDK

Trace workflows, search tools, and export telemetry from your application.

CLI

Query traces, run SQL, and operate the control plane from the terminal.

Onboarding

1

Sign in to the console

Open your Caveman Cloud console and sign in. If you were invited to a project, you land directly in that workspace.
2

Create a project and API key

Go to Gateway, Connect to create a project, then generate a Cave API key. Copy the key and your gateway URL. The gateway URL has no trailing slash.
3

Set environment variables

Export CAVE_API_KEY with your Cave key and CAVE_GATEWAY_URL with your gateway origin. Keep your upstream provider key separate.
4

Swap the base URL

Change the baseURL on your OpenAI or Anthropic client to ${CAVE_GATEWAY_URL}/openai/v1 or the matching provider path. Add the x-cave-upstream-key header when the provider key is not stored in Caveman Cloud.
5

Send a request and view the trace

Run one chat completion or agent turn. Open Traces in the console to inspect the request, model, usage, and latency.

Core surfaces

Savings evidence

Caveman Cloud reports three kinds of numbers, and it keeps them separate:
  • Measured spend records observed usage and its pricing basis.
  • Inferred savings describe an estimate or counterfactual, not a guaranteed invoice reduction.
  • Verified savings require a supported accounting method and qualifying evidence. They stay zero when that evidence is absent.
See Savings Evidence for the full contract, Optimization Catalog for what each optimization changes, and Reporting Savings for turning these numbers into a monthly finance report.