Skip to main content
The @caveman-ai/sdk package is a zero-runtime-dependency TypeScript client for Caveman Cloud. It wraps the gateway’s /sdk/v1/* contract, so you can trace agent workflows, delegate compression, search tool catalogs, proxy provider calls, and export OpenTelemetry spans without adding extra dependencies to your project. This page covers installation, configuration, per-request controls, receipt parsing, Cloud control-plane resources, and error handling for the TypeScript SDK.

Install

Install the package from npm. It ships as an ES module and requires Node.js 22.13 or newer.
Import the Cave client and the types you need:

Configure the client

Create a Cave instance with your API key, gateway base URL, and an agent slug. The agent slug labels every span and trace the SDK records. Optionally set a defaultWorkflow so every traced unit of work is tagged to a workflow unless you override it per call.
The constructor throws if apiKey, baseURL, or agent is missing. The baseURL should not have a trailing slash. For example, if your gateway URL is https://gateway.example.com, OpenAI calls go to https://gateway.example.com/openai/v1.

Trace workflows

Use cave.trace() to wrap a unit of work. Inside the callback you get a CaveTrace with helpers for tools, model calls, artifacts, and checkpoints. The SDK POSTs telemetry to the gateway and swallows transport failures so they never break your agent.

Checkpoints

Offload conversation context to the gateway and retrieve it later with a reversible source_ref.

Context packing

Send caller-owned context fragments to the gateway and get back the subset that fits under a token budget, plus the exact IDs of everything deferred.
Packing is a lossy selector, not a compressor. basis is always "inferred" and omitted items are named in deferredIds so your application can re-supply them.

Artifacts

Page a large payload to gateway storage and return a compact stub the model can expand on demand.

Per-request control

Optimizations are a project-level setting, but a single request can ask for or disable specific optimizations. The gateway checks the project policy, capability, and evidence gate before applying anything; a request never grants itself permission.
Unknown tokens are rejected with 400 cave_invalid_optimize_override. Read the response receipt to see whether an optimization was applied or denied.

Model routing

Send model: "cave-auto" and let the project’s router pick the model on every request. A named model is kept as-is. If the project has no provider connection or baseline model configured, the gateway returns 400 cave_routing_not_enabled.

Receipts

Every gateway response carries per-call disclosure headers. Use parseReceipt to read them. Headers the gateway did not send stay absent, never defaulted.
tokensBefore and tokensAfter are inferred estimates from the compressor, not provider counts. If recoveryHandle is present, you can retrieve the byte-exact original.

Provider clients

cave.openai(), cave.anthropic(), cave.gemini(), and cave.vertex() return thin clients proxied through the gateway. Each has a .raw escape hatch for native provider paths.
You can attach a latencyClass hint ("interactive", "background", or "offline") to any request. Anything other than "interactive" sets the x-cave-async header so the gateway can defer it.

Compression

cave.compress() POSTs a payload to the gateway and returns the Engine report. The SDK delegates all compression decisions; it never reimplements a compressor client-side.
If anything goes wrong during transport or parsing, the SDK passes the original through unchanged: ratio is 0, there is no recoveryHandle, and basis is "inferred". cave.tools() returns a handle with the full catalog, a deferred initial subset, and a search() method that queries the gateway.
search() is async and must be awaited. The gateway honors "embeddings" only when it has an embedding provider wired. cave.toolSearch() is the flat variant when you manage the catalog yourself.

The OTel exporter

cave.exporter() returns an OTelExporter that builds OTLP/JSON by hand and POSTs to the gateway’s /v1/traces endpoint. No external OpenTelemetry SDK or collector is required.

Retry-loop breaker

Interrupt an agent that is stuck re-issuing the same tool call.

Shared context

Hand context between agents in the same project. Keys are tenant-scoped, so a peer in another project cannot read them.

Prompt snippets

Build deterministic internal brevity instructions for system prompts.

Cloud control-plane resources

Import Cloud from @caveman-ai/sdk/cloud to work with Caveman Cloud control-plane resources. Use a Cloud access token, separate from gateway inference keys.
Receipted writes require an explicit idempotencyKey. The client does not automatically retry writes. Accepted work is not a passing result: always read the persisted result or use the CLI wait workflows.

Error handling

The SDK throws descriptive errors for common gateway and local failure modes: On network or parse errors, compression and other delegated operations fall back safely. The SDK never rewrites bytes itself and never reports a saving the engine did not hand back.

Examples

Export OTel spans

The bundled exporter example records a chat completion span and a child tool span, then ships them to the gateway.

Next steps

  • Python SDK for the same gateway contract in Python
  • CLI for running workloads, managing resources, and CI workflows
  • Gateway for provider integration guides