Replay a decision
The selected-step lab compares original and edited context across repeated model responses. Repeat rates describe observed behavior; they do not recover private reasoning.
OBSERVABILITY FOR CODING AGENTS
Bring prompts, tool calls, reported results, and costs into one local dashboard. RayTrace CLI is now open source. Install it, inspect the code, and make it yours.
“Fix the failing authentication test.”
src/auth.ts · tests/auth.test.ts
Bash
gVisor exec checkpoint · process exit
RayTrace CLI is live. The local recorder and dashboard, free to use and MIT licensed.
Explore the code on GitHub ↗What did it see? / What did it try? / What actually ran? / What would change?
01 / INSPECT
A final answer only tells part of the story. RayTrace connects prompts, context, tool calls, and runtime evidence so you can inspect the steps that led there.
Inspect the input snapshot for a model request: the prompt, tool definitions, prior results, and file excerpts included in its context.
// Excerpt included in the model request
test('session has not expired', () => {
expect(session.expiresAt)
.toBeGreaterThan(Date.now());
});A file in the repository is not necessarily a file the model saw.
02 / EXPERIMENT
Was a piece of context useful? Select a step, edit or remove what the agent saw, and compare the continuation with the original run.
The sandbox continuation below is an advanced development workflow. The CLI release focuses on local capture and inspection. View release scope →
The selected-step lab compares original and edited context across repeated model responses. Repeat rates describe observed behavior; they do not recover private reasoning.
For eligible sandboxed Claude Code sessions, Playground restores the project before a step and runs a fork. Compare file changes, steps, and an optional check command.
Models can vary even when nothing changes. Playground can interleave edited runs with unchanged runs, so you can examine that variation alongside your experiment.
KNOW THE BOUNDARY Forks need a recorded sandbox session with snapshots. Decision replays and Codex continuations have different setup and provider requirements.
03 / UNDER THE HOOD
The CLI records Claude Code transcripts through a hook and Codex model exchanges through a local proxy. A bundled dashboard reads your local trace history.
* OpenRouter is required for the CLI’s Codex launcher and optional generated summaries. Basic Claude Code recording uses a local hook and its usual model connection.
Trace data stays in a local SQLite database. Content-addressed payloads reuse repeated context instead of storing the same history for every request.
Outside the packaged CLI, a gVisor sandbox inside a Lima Linux VM can independently collect process-start and exit events. Read the sandbox requirements →
A process starting or exiting successfully does not prove the task was done correctly. Missing evidence means unknown. RayTrace does not reconstruct private chain of thought.
04 / GET STARTED
Install the CLI, set up Claude Code capture, and inspect your first session. The recorder and dashboard start together.
Built for experimentation. RayTrace is an early prototype for local debugging and evaluation. Captured prompts and file contents can be sensitive; the local services do not yet provide authentication or encryption at rest.
npm install -g @raytrace-cli/cliraytrace setupraytrace opennpm install -g @raytrace-cli/cli
raytrace setup
raytrace openAdd your OpenRouter key during setup, then run raytrace codex in your project. Model requests use OpenRouter and provider billing.