product

The flight recorder for AI agents, in four parts.

Route every call, see what your agents did, prove it to a third party, and ship changes you can defend. 34 capabilities you can use today.

What the flight recorder does

Recorded before it is dispatched

Every call the hosted gateway admits is written to the audit ledger before it goes to your provider. If the ledger cannot take the write, the call does not go out.

How the ledger works →

Verifiable without trusting us

SHA-256 hash-chained, Ed25519-signed and anchored to Sigstore's public log. Three open verifiers, in Rust, Python and TypeScript, check it offline.

Security posture →

Every attempt, and the model that answered

Each retry and failover hop is on the trace with its status and timing, and every call records the model you asked for next to the one the provider says served it.

What a trace shows →

2 ms added at p50, on your own keys

Measured with LiteLLM's open AIGatewayBench, full table published. 205 providers through one endpoint, your keys, 0% markup.

The benchmark, method and caveats →

gateway

One endpoint for every model you use.

Point your OpenAI or Anthropic client at Tracelane. Your calls go to your provider with your own key, every one of them is recorded, and the routing is yours to change without a deploy.

Control — one endpoint that routes, caps and swaps models, changed in settings, not in code.

Gateway docs →

2 / 4 / 5 ms added at p50 / p95 / p99, under steady traffic AIGatewayBench, 2026-09-07, two runs on a 4-vCPU host

8 capabilities · 3 proven on production

OpenAI-compatible endpoint

Change one base URL; any OpenAI-compatible client works, and every call is traced.

Anthropic-native endpoint

proven

Point Claude Code or any Anthropic client straight at the gateway and get full tracing, tokens and cost.

Your keys, 0% markup

Store provider keys encrypted per workspace, check on demand that a key still works, and pay your provider directly.

Name your own models

proven

Call a model by a name you control, like "fast", and repoint it in settings — the trace records both names.

Failover you configure

Turn on cross-provider failover for your whole workspace and choose the fallback models, in order.

Exact-match response cache

proven

Serve an identical repeated request from cache; a response header says it was cached.

Guardrails at the gateway

Guardrail rails check requests at the gateway — cost and step caps, secrets and personal data, tool pinning, format and topic — and flag or redact what matches. Which rails you get depends on your plan.

Breakers and a real cost per call

Each provider gets its own circuit breaker, and every call carries its USD cost from a model price catalogue.

observability

See what your agents actually did.

Every call, tool and retry, readable as the conversation it was — searchable, shareable, and in the OpenTelemetry format you already use.

Visibility — when an agent misbehaves, the answer is already recorded.

Observability docs →

OTel native — OTLP in, GenAI conventions send OTLP directly; spans follow the OpenTelemetry GenAI semantic conventions

10 capabilities · 9 proven on production

Search and slice traces

proven

Search by span name or content, filter, group and sort, and export what you see to CSV or JSON.

Traces read as conversations

Read a trace as the transcript it was, with the span tree and an inspector beside it.

One lane per agent

proven

See multi-agent runs with a lane per sub-agent and a failures-only filter.

Every retry and failover hop

proven

See each attempt a request made — retries and failover hops with status and timing — not only the one that worked.

Calls that went wrong, flagged

proven

A trace is marked when the model that answered is not the one you asked for, the reply came back empty, or its stream was cancelled, and when an OpenAI- or Anthropic-style provider or your OpenTelemetry spans report a token-limit or content-filter stop. Filter the list to just those.

Label every request

proven

Tag calls with an environment, release, service and your own tags, through the gateway or over OpenTelemetry; every span keeps them.

Know every calling agent

proven

See which agents and clients call your gateway, each with its own profile and history.

Traces by your end user

proven

Send your own user id and filter sessions and traces by the person who started them.

Share one trace

proven

Send a revocable, account-free link to a single trace, with its ledger badge.

Dashboards, SLOs and alerts

proven

Build your own dashboards, read latency and error SLOs, and route alerts to Slack, Discord or a webhook.

audit ledger

A record your auditor can check without us.

Every call through the gateway, and every policy result on it, goes onto a tamper-evident hash chain for your workspace, signed with your workspace's own key and anchored to a public transparency log. Spans you send directly over OTLP or an SDK are recorded, not chained.

Proof — evidence a third party can verify offline, without trusting us.

Audit ledger docs →

3 independent offline verifiers Rust, Python and TypeScript — run without network access

8 capabilities · 1 proven on production

Tamper-evident ledger

Every gateway call and policy result is chained, so a later edit, deletion or reorder is detectable.

Signed with your workspace's key

Batches are signed with a key that belongs to your workspace, not one shared across customers.

Anchored to a public log

Batch roots are anchored in the public Sigstore Rekor log, which anyone can look up.

Verify offline, three languages

Verify an exported bundle with no network access, in Rust, Python or TypeScript.

A verifier an auditor runs

Hand an auditor one signed binary that exits non-zero when the chain does not hold.

Verify on any plan

Check your chain in the product at no cost, including a recompute in your own browser.

Written in one transaction

proven

Each ledger row is committed in the same database transaction as the chain it extends, and archived hourly.

EU AI Act Article 12 export (Enterprise)

On Enterprise, export the record-keeping pack mapped obligation by obligation; an export that cannot complete fails loudly instead of handing you half. Per-record inclusion proofs and a signed completeness attestation are on the roadmap.

evals

Ship prompt changes you can defend.

Turn production traces into test cases, score live traffic, and let CI refuse a change that falls below the floor you set.

Quality — prove a prompt or model change is better before it reaches users.

Evals docs →

CI gate on your own floor the eval gate fails a build when quality falls below a threshold you set

8 capabilities · 5 proven on production

A CI eval gate

proven

Fail a pull request when a dataset's score falls below your floor — with separate exit codes for "worse" and "couldn't measure".

Score live traffic

proven

Let an LLM judge score a sampled share of production traffic, under a spend cap you set.

Canary a prompt version

proven

Send a new prompt version to a share of your users, each kept on the same version, before promoting it.

Versioned prompts, signed changes

Version prompts and promote or roll back; every change is a signed record in your ledger.

Review queues

proven

Send failing traces to a queue, write the right answer, and keep it as a test case.

Datasets from production

Keep datasets that persist, and add a trace's input from the trace page in one click.

Compare against a baseline

Compare an experiment against a baseline case by case, not only on the average.

Playground to trace

proven

Try a prompt against your connected provider and land on the trace it produced.

Product tour

Watch the recorder in action.

A short walkthrough of traces, gateway decisions and audit evidence.

Start free Read the docs

Everything shipped, dated and labelled: the changelog.