Skip to main content
The caveman-sdk distribution provides a stdlib-only Python client for Caveman Cloud. The import package is caveman_cloud. It uses only the standard library (no requests, no httpx) and wraps the gateway’s /sdk/v1/* contract so you can trace agent workflows, delegate compression, search tool catalogs, proxy provider calls, and export OpenTelemetry spans without adding third-party dependencies. This page covers installation, configuration, per-request controls, receipt parsing, Cloud control-plane resources, and error handling for the Python SDK.

Install

Install the package from PyPI. It requires Python 3.13 or newer.
Import the Cave client and the types you need:
The import package is always caveman_cloud, regardless of how you install it. The distribution name on PyPI is caveman-sdk.

Configure the client

Create a Cave instance with your API key, gateway base URL, and an agent slug. The agent slug labels every span and trace the SDK records. Optionally set a default_workflow so every traced unit of work is tagged to a workflow unless you override it per call.
The constructor raises if api_key, base_url, or agent is missing. The base_url should not have a trailing slash. For example, if your gateway URL is https://gateway.example.com, OpenAI calls go to https://gateway.example.com/openai/v1.

Trace workflows

Use cave.trace() as a context manager to wrap a unit of work. It yields a Trace with helpers for tools, model calls, artifacts, and checkpoints. The SDK POSTs telemetry to the gateway and swallows transport failures so they never break your agent.

Checkpoints

Offload conversation context to the gateway and retrieve it later with a reversible source_ref.

Context packing

Send caller-owned context fragments to the gateway and get back the subset that fits under a token budget, plus the exact IDs of everything deferred.
Packing is a lossy selector, not a compressor. basis is always "inferred" and omitted items are named in deferred_ids so your application can re-supply them.

Artifacts

Page a large payload to gateway storage and return a compact stub the model can expand on demand.

Per-request control

Optimizations are a project-level setting, but a single request can ask for or disable specific optimizations. The gateway checks the project policy, capability, and evidence gate before applying anything; a request never grants itself permission.
Unknown tokens are rejected with 400 cave_invalid_optimize_override. Read the response receipt to see whether an optimization was applied or denied.

Model routing

Send model="cave-auto" and let the project’s router pick the model on every request. A named model is kept as-is. If the project has no provider connection or baseline model configured, the gateway returns 400 cave_routing_not_enabled.

Receipts

Every gateway response carries per-call disclosure headers. Use parse_receipt to read them. Headers the gateway did not send stay None, never defaulted.
tokens_before and tokens_after are inferred estimates from the compressor, not provider counts. If recovery_handle is present, you can retrieve the byte-exact original.

Provider clients

cave.openai(), cave.anthropic(), cave.gemini(), and cave.vertex() return Provider objects proxied through the gateway. Each exposes a .raw(path, body) escape hatch for native provider paths.
You can pass a latency_class ("interactive", "background", or "offline") to any request. Anything other than "interactive" sets the x-cave-async header so the gateway can defer it.

Compression

cave.compress() POSTs a payload to the gateway and returns the Engine report. The SDK delegates all compression decisions; it never reimplements a compressor client-side.
If anything goes wrong during transport or parsing, the SDK passes the original through unchanged: ratio is 0.0, recovery_handle is None, and basis is "inferred". cave.tools() returns a handle with the full catalog, a deferred initial subset, and a .search() method that queries the gateway.
The gateway honors "embeddings" only when it has an embedding provider wired. cave.tool_search() is the flat variant when you manage the catalog yourself.

The OTel exporter

cave.exporter() returns an OTelExporter that builds OTLP/JSON by hand and POSTs to the gateway’s /v1/traces endpoint. No external OpenTelemetry SDK or collector is required.

Retry-loop breaker

Interrupt an agent that is stuck re-issuing the same tool call.

Shared context

Hand context between agents in the same project. Keys are tenant-scoped, so a peer in another project cannot read them.

Prompt snippets

Build deterministic internal brevity instructions for system prompts.

Cloud control-plane resources

Import Cloud from caveman_cloud to work with Caveman Cloud control-plane resources. Use a Cloud access token, separate from gateway inference keys.
Receipted writes require an explicit idempotency_key. The client does not automatically retry writes. Accepted work is not a passing result: always read the persisted result or use the CLI wait workflows.

Error handling

The SDK raises descriptive errors for common gateway and local failure modes: On network or parse errors, compression and other delegated operations fall back safely. The SDK never rewrites bytes itself and never reports a saving the engine did not hand back.

Examples

Export OTel spans

The bundled exporter example records a chat completion span and a child tool span, then ships them to the gateway.

Next steps

  • TypeScript SDK for the same gateway contract in TypeScript
  • CLI for running workloads, managing resources, and CI workflows
  • Gateway for provider integration guides