跳到正文
Hacker News · AI· rganupa·· 3 小时前AI 评分37

Rashomon:为 AI 编程智能体提供独立执行记录,仅支持 Claude Code

Show HN: Rashomon – An independent execution record for AI coding agents

AI 导读

Rashomon 是一款为 Claude Code 提供独立执行记录的 Alpha 工具(v1.1.0),仅支持 macOS 和 Linux。它记录智能体每一步命令与文件编辑及其成败,并与智能体自述对照,若失败步骤未被总结提及则在回合结束时提示,报告还会列出主对话不显示的子智能体操作。该工具不存储提示词、文件内容或工具输出,无账号、无遥测、无网络代码。

正文

Your coding agent can tell you it finished the job. Did it actually?

rashomon keeps an independent record of what your agent actually did, then puts that record beside the agent's own account.

Join the Rashomon Discord

For Claude Code only, on macOS and Linux. Alpha: v1.1.0.

Claude says all tests pass; rashomon's end-of-turn line reports a recorded failure, and an excerpt of the report shows it

In plain words

When Claude Code finishes a task, it tells you what it did: "Done, all tests pass." That summary comes from the agent itself.

rashomon keeps its own record of every step the agent takes, like each command it runs and each file it edits, and whether each one worked. It puts that record beside the agent's summary: if a step failed and the closing message uses none of rashomon's failure words, it tells you at the end of the turn. Every report also lists what subagents did, which the main conversation does not show.

You keep working the way you do now. rashomon stays quiet unless something is worth a look. The name comes from Rashomon, the film in which witnesses give different accounts of the same event.

Get started

Pick one of these two ways. Using both records everything twice.

As a Claude Code plugin:

claude plugin marketplace add altrace-dev-role/rashomon
claude plugin install rashomon@rashomon
claude plugin enable rashomon@rashomon

Then start a new Claude Code session. Inside it, /rashomon:report shows what was recorded.

From the command line (macOS and Linux):

curl -fsSL https://raw.githubusercontent.com/altrace-dev-role/rashomon/main/install.sh | sh
rashomon watch

Then use claude as usual, and run rashomon report whenever you want to look. The installer puts rashomon in ~/.local/bin; if your shell cannot find it, add that folder to your PATH.

What you'll see

Most of the time, nothing: a turn with nothing worth a look prints nothing. When something is off, one line appears at the end of the turn (or, if you refused a permission prompt, when you send your next prompt, marked previous turn):

※ rashomon: 1 recorded failure.
            → rashomon report --session <id>

With the plugin, the arrow points at /rashomon:report --session <id> instead. That command shows the whole story. Here is an example.

What a disagreement looks like

Here the tests failed while a helper agent (a subagent) looked through the code, yet the agent's closing message says the tests pass. The session is built by test/acceptance/readme_test.go, which feeds hook payloads through the real recorder, and that test fails if this excerpt and the render ever differ. This excerpt is rashomon report exactly as it renders that session:

  the agent's account:
    "Done. I refactored ParseConfig and all tests pass."
  subagents: 1
    agent-a41f (Explore): 3 declarations, 3 executions, 1 Bash
    (these calls do not appear in the main transcript)
  failed calls: 1
    the final message contains none of these 43 words: fail, failed, failing, error, errors, couldn't, could not, unable, not able, didn't, did not, blocked and 31 more of 43 (--json lists them all)
  test runs: 1 (0 ok, 1 failed)

The first line quotes the agent. The rest comes from rashomon's own record, checked against those words: one call failed (its record holds exit code 1, and --chain shows which call it was), and the subagent's three calls are recorded under the subagent's own transcript, not the main one. The last line counts the calls rashomon recognised as a test runner (go test, pytest, npm test and a fixed list of others) by how they ended. The report lists the failure words the summary does not use; it does not guess why.

What a report tells you

  • What subagents did: the helper agents Claude starts on its own, whose steps the main conversation never shows.
  • Which steps failed, and whether the agent's closing summary uses any of rashomon's failure words. It lists the ones that are absent; it never guesses at intent.
  • What ran differently from what was asked, for example another hook that rewrote a command's input before it ran.
  • What rashomon could not see, on every report, even a healthy one.

Try it yourself

You can produce the same kind of report in a few minutes:

  1. Install rashomon (see Get started).
  2. In a project you have open, break one test on purpose.
  3. Start claude. Ask it to have a subagent look through the code, then to run the tests and finish with a one-line summary.
  4. Run rashomon report (or /rashomon:report inside Claude Code).

The report quotes the summary above the subagents and failed calls lines. Whether the failure is flagged depends on what the agent writes. If the summary contains none of rashomon's 43 failure words (matched as substrings, so "debug" counts as "bug"), the failed calls line lists them. Most honest summaries contain one, and then the report says so instead.

Common questions

Does it see my code or my prompts? It sees them in passing and keeps none of them. Claude Code hands every hook the tool call's full input (for Write and Edit, that includes the file's text) and hands the next-prompt hook your prompt; rashomon reads what it needs and discards the rest. It stores no prompts, responses, file contents or tool output. A command is reduced to its program name, an argument count and a keyed digest, by a parser that names nothing when it is unsure. The one thing it keeps from inside a command is any hostname the command names (and the host of a WebFetch URL), stored in clear so the report can list it. The working directory and transcript path are stored too; see What is recorded. When you run rashomon report, it quotes the agent's final message from Claude Code's own transcript; that quote is never stored.

Does it send anything anywhere? It has no account, no telemetry and no network code of its own, and what it records is kept in a folder on your machine. No rashomon package imports a networking package, and CI checks that on every pull request, every push to main and every release tag. On Linux, CI also runs the declaration hook with the network switched off; the other hooks, the report and the end-of-turn line have no such runtime test yet.

Can it break my agent? No. It never stops a command, and if something inside it goes wrong, it exits quietly so Claude Code carries on.

Does it work with Cursor or Codex? No. This release supports Claude Code only. Cursor (by default) and Codex (after /import) read Claude Code's settings, so they can trigger the recorder anyway, and what it records for them can be wrong. If you use Cursor, see Scope for the setting to turn off.

Is it a sandbox or a security tool? No. It does not stop or contain the agent, and it cannot stop a determined agent from changing its records. It is a second, independent account of what happened, to set beside the agent's own.

Can I see network activity? rashomon itself observes no network traffic in this alpha: the report says destinations: not observed in this alpha, and --chain lists only the hosts each command named. If you run Claude Code inside the nono sandbox, rashomon report --nono-audit <path to nono's audit-events.ndjson> adds what nono allowed and refused in that session's window. This is experimental, and nono writes that file when its session ends. See Inside a sandbox first.

How do I turn it off? Plugin: open /plugin in Claude Code and turn it off. Command line: rashomon detach. Your records stay in ~/.local/state/rashomon (by default) until you delete that folder.

Is it finished? No. This is an alpha: Windows is not supported yet, and network destinations are not observed in this release.

The details

Everything below is the precise version: what gets installed, what is stored, and what the report can and cannot see.

Install options, in detail

As a Claude Code plugin (macOS and Linux):

claude plugin marketplace add altrace-dev-role/rashomon
claude plugin install rashomon@rashomon
claude plugin enable rashomon@rashomon

Installing does not start recording; enabling does. Start a new Claude Code session after enabling. /plugin shows it and turns it off again, and /rashomon:status / /rashomon:report work inside Claude Code. The plugin installs the release's rashomon-plugin.zip, which carries a prebuilt recorder per platform, because a plugin cannot compile Go at install time. So it needs a tagged release to exist. To try the plugin from a checkout instead:

scripts/build-plugin.sh
claude --plugin-dir ./plugin --settings '{"enabledPlugins": {"rashomon@inline": true}}'

The plugin ships disabled, and a plugin loaded with --plugin-dir (Claude Code names it rashomon@inline) is enabled only through settings. Without the --settings flag the session loads the plugin and records nothing.

If you already ran watch, run rashomon detach first. With both installed, every call is recorded twice, and the report marks the session duplicate_declarations. rashomon status names the overlap for a plugin installed from the marketplace; it cannot see one loaded with --plugin-dir.

As a settings install (Windows is not usable yet; see Status). From a release, on macOS or Linux:

curl -fsSL https://raw.githubusercontent.com/altrace-dev-role/rashomon/main/install.sh | sh

It downloads the release archive for your platform, checks it against the release's checksums.txt, refuses to install on a mismatch, and puts one binary in ~/.local/bin (set RASHOMON_INSTALL_DIR to change that). Or with Go:

go install github.com/altrace-dev-role/rashomon/cmd/rashomon@latest

This writes to $GOBIN, or $(go env GOPATH)/bin when GOBIN is unset; make sure that folder is on your PATH. A binary built this way, or from a clone, reports its version as dev: only release builds carry the version number. Or clone and build:

git clone https://github.com/altrace-dev-role/rashomon
cd rashomon && go build ./cmd/rashomon

Install to a location that will not move. watch writes the running binary's absolute path into your Claude Code settings, and Claude Code executes that path on every tool call. If you built inside a clone you later delete, every tool call fires a hook that cannot start. Build into a directory you keep (or go install it), and re-run watch after any move. watch records the path with symlinks resolved, so a stable symlink that points into a clone still breaks when the clone goes. (watch refuses a binary whose path runs through a go-build directory, which is where go run builds it.)

Why a broken hook matters more than usual: a PreToolUse hook that exits non-zero in the wrong way can block the tool call, and the user sees Claude Code failing rather than this program. Every recoverable failure here is engineered to exit 0 for that reason (a Go runtime fatal error cannot be caught, so those are prevented structurally instead, for example by capping how much input a hook reads); see docs/design-notes.md.

What watch sets up

rashomon watch                      # install the recorders (the only command that installs anything)
claude                              # work normally
rashomon report                     # read the sessions back

watch adds eight entries to your Claude Code settings: PreToolUse, PostToolUse, PostToolUseFailure, SessionStart, SessionEnd, Stop, StopFailure, UserPromptSubmit. The first three match * (every tool); the rest carry no matcher, so they fire on every occurrence of their event. All eight run a 5-second timeout except the three recap entries, which get 10: finding a turn's boundaries means reading a whole run, not one call. rashomon detach removes them again and leaves every other entry's value byte-for-byte as found. The top-level keys and the hooks block are re-written with two-space indentation, so a file kept in another layout comes back with mixed indentation.

Stop and StopFailure drive an exception-only recap: after a turn, it prints at most one line, and only when there is something worth looking at (a recorded failure that the agent's final message does not acknowledge, meaning it uses none of the failure words; a declaration without recorded execution; coverage that did not verify; a truncated/unknown projection; a failed test command that passed after the only recorded edits were to files named like tests; or the same test command passing and failing with no recorded file edit between). A file edit there is any recorded call that may change files, not only an Edit or a shell rm: a git checkout, an npm install, a sed -i, an MCP tool, another test command (jest -u rewrites snapshots) all count, and only reads, web fetches, subagent launches and Claude Code's own tools that write no source or test file (TodoWrite, TaskCreate, TaskUpdate, TaskList, TaskGet, TaskOutput, TaskStop, AskUserQuestion, EnterPlanMode, ExitPlanMode, BashOutput, KillShell, KillBash, Skill, ToolSearch, SendMessage, CronCreate, CronDelete, CronList, ListMcpResourcesTool, ReadMcpResourceTool) do not. A shell read or fetch counts as well when its line may write, which the record keeps as one bit (may_write): an output redirect to a file, a download (curl -o/-O, with the file attached or not, as in curl -o./calc.go, and wget), a command or process substitution ($(…), backticks, <(…), >(…)), find -delete/-exec/-execdir, xargs, tee, rsync or scp anywhere on the line, or a later pipeline or list stage whose program is not a read (grep -rl … | xargs sed -i). That under-claims by design. It can still miss a change: a read or fetch that writes through an option not on that list (find -fprint f, curl -D f, curl -c f), or anything done outside the session's own calls. One known gap: Claude Code discards what a StopFailure hook prints, so a turn that ends in an API error shows no line, and in this release the next prompt does not show it either. The line points to rashomon report --session <id> for the detail. A clean turn prints nothing at all; rashomon status says whether a turn has actually been evaluated, so silence never gets read as proof the turn was clean.

The two test-bending lines have limits of their own:

  • A test run is a shell call whose whole command line is one of these runners, optionally after cd DIR && or NAME=value assignments: pytest, python -m pytest, python3 -m pytest, jest, vitest, mocha, rspec, phpunit, ctest, tox, nox, go test, cargo test, npm test, npm t, npm run test, yarn test, pnpm test, bun test, dotnet test, mvn test, mvnw test, gradle test, gradlew test, make test, npx jest, npx vitest, uv run pytest, poetry run pytest and bundle exec rspec (a program is matched by its base name, so ./gradlew test and ./mvnw test count). Each may follow timeout N and then time (timeout 120 go test ./..., time pytest), which pass the runner's exit status through; when the timeout fires, its own status 124 is read as no result, neither passed nor failed. These wrappers are counted because their exit status is the runner's.
  • A runner followed by a pipe or a list (| tail, 2>&1 | grep, && echo ok, ; echo done) is not counted; a redirection alone (2>&1, > out.txt) is. Without pipefail a pipe's status is its last program's, and after && or ; the line's success is the next command's. So piped runs such as go test ./... 2>&1 | tail -20, which are much of what Claude Code writes, are invisible to both patterns, and a session whose tests ran only that way shows no test runs block at all.
  • Two runs are the same command only when their command lines are identical character for character. go test ./... with two spaces, a trailing space, a cd /repo && prefix or a CGO_ENABLED=0 prefix is another command, and never pairs with the plain form.
  • A runner is on the list when it is a known test tool, a build tool or launcher given its test command (go test, npm t, npm run test, python -m pytest), or a listed wrapper that passes its runner's exit status through, and that does not keep lint out: go test runs vet, npm test runs a pretest script, and tox's default envlist or a make test target can include lint. A lint failure fixed only in a file named like a test then reads as the tests-only pattern.
  • Runs pair only within one directory: each call's record carries a keyed digest (never the path) of the directory its command starts in: the reported cwd, or where its leading plain cd DIR && steps lead (cd /repo/web && go test ./... is keyed on /repo/web wherever the shell was). Two runs of one command line in two directories are not the same run. So a subagent's go test ./... pairs with the main agent's only when both ran it from the same directory, and a repeated relative cd sub && go test ./..., whose second call starts in sub and so targets sub/sub, does not pair with the first.
  • cd DIR && go test ./... is a test run, so a failed cd counts as a failed test run. Leaving cd … && out would lose most real runs.
  • Claude Code moves a command to the background when it reaches its timeout (two minutes by default) or when you press Ctrl+B, and its PostToolUse then fires before any test has finished. The execution record says so (backgrounded), as it does for a run_in_background launch (a Ctrl+B background is recognised only if Claude Code marks it with backgroundTaskId or backgroundedByUser, which was not measured), and such a run is read as outcome unobserved: it is neither ok nor failed in test runs, completes no pattern, and sits under unknown in --timeline, where it is never offered as a later success. Records written before schema 3 cannot say, and a run backgrounded there still reads as ok.
  • A test edit is an Edit, Write or NotebookEdit that ran ok on a file named like a test (test-file). Deleting or moving a test through the shell, or regenerating snapshots or golden files (jest -u, a -update flag), is a shell call with no label, so it never completes the tests-only pattern.
  • No pair is reported for a session holding a call whose declaration was lost (the store has its execution record or terminal, and no declaration): that call may have changed a file between two runs, and nothing records where it fell. The test runs block then says no pair is looked for, and how many such calls there are.
  • The end-of-turn line names only a pair whose two runs are both in its turn, and reads no denied set, so a pair across two turns, or one completed only across a denied edit, can show in the report and not in the line. The other way round, a pair the line named is missing from a later report when a call whose declaration was lost is recorded after it, in that turn or a later one.

report --json carries the same facts: each session's test_runs (null for a session with no schema 3 declaration: records that predate schema 3, or no tool calls at all; otherwise only runs that ended ok or failed are counted; interrupted, denied, backgrounded, timed-out and unrecorded runs are not, and a session that spans the upgrade is counted from its first schema 3 call; then ok, failed, undeclared, how many calls lost their declaration (when it is not 0, no pair was looked for), tests_only_then_green as [earlier, later] seq pairs, and flaky as {"seqs": [earlier, later], "first_failed": true|false}, where earlier and later are declaration order, the order the runs started, which overlapping runs in parallel agents may not have finished in), and a test_bending (kind, since_seq) on the timeline row that completes a pair, null on every other row.

Saying No at a permission prompt interrupts the turn, and Claude Code fires no Stop after an interrupt. UserPromptSubmit catches that case: when you send your next prompt it checks the turn that just ended, and if no recap reached it and it has something to show, prints the line then, marked previous turn. The catch-up has no final message to check, so it can show every finding except an unacknowledged failure.

What is recorded, and what never is

Recorded: identifiers, shapes and hostnames. tool_use_id, session_id, prompt_id, agent_id, agent_type, transcript_path, cwd, permission_mode, tool_name; the call's program, verb_class, argument count, and a keyed digest (HMAC) of its full input: the whole command line for Bash, the whole tool input for other tools; for Bash, one bit (may_write) saying the line may write files whatever its program is; for every call, a keyed digest of the directory its command starts in: the reported cwd, or where its leading plain cd DIR && steps lead (cwd_digest); how it ended (outcome, exit_code, is_interrupt, duration_ms, and backgrounded: whether a Bash call's result arrived while it was still running in the background); and a file_label classifying the path named by a Read, Edit, Write or NotebookEdit call into categories such as ssh-key, env-file, cloud-config, credential-shaped, certificate or test-file (paths touched from Bash are not labelled). A shell call whose command is a recognised test runner has verb_class test, unless an argument on that runner's fixed refusal list makes it do something else (such as compile, list, dry-run, skip the tests, watch, help, version, or a named tox -e/nox -s target), or the line sets PYTEST_ADDOPTS before pytest. The refusal lists name particular spellings and are not complete: a combined short flag, an option set in a config file or by an earlier call's environment, or a spelling not on the list is still counted as a test run. The arguments that decided it are compared against the lists and not kept.

Note that cwd and transcript_path are filesystem paths and carry directory names. docs/store-schema.json is the field list for the record files; a record carries no key the schema does not declare (the one open object is rule_match).

Hostnames are stored in clear, because the report has to name them: report --chain lists the hosts each call named beside that call, and a digest cannot be rendered back into a name. They are extracted from WebFetch.url, from any http(s), ws(s), ssh://, git+ssh:// or git@host: address in a Bash command line, and from the destination arguments of ssh, scp, sftp and rsync, whether or not the call reached it. Only those two tools are read.

Never recorded: prompts, responses, argument values (other than those hostnames), file contents, tool output. Not redacted: structurally absent from the store. The recorder's payload types declare no field for tool output, so it is never decoded into a value and no record has a field that could hold it. The one thing read from a Bash call's tool_response is whether its keys backgroundTaskId and backgroundedByUser are present, as one bit; the task id and every other key (stdout, stderr) are not decoded. The raw hook input is read into memory before decoding, and a failed call's error message, which can quote command output, is decoded and reduced to an exit code; nothing else of it is kept. Commands are reduced to a program name and an argument count, by a parser that names nothing when it is unsure. It can still misread a rare shape; how it decides is described in #29. A misread can store a fragment of a command, so report one privately (see Contributing).

Those rules govern the store. The report additionally reads Claude Code's own transcript at render time: it quotes the agent's final message, and checks tool results for the message Claude Code writes when you refuse a call. That is content: none of it is written to the store or transmitted, and --redact drops the quote entirely rather than partially cleaning prose that may name anything.

The digest is an HMAC under a random per-install key kept in the store (install.key). The same command digests differently on two machines, so digests cannot be matched against a dictionary built elsewhere. Anyone holding the store holds the key, though, and can test guessed commands against it. Store directory 0700, files 0600.

Where the store lives, and how to remove it

$RASHOMON_HOME, else $XDG_STATE_HOME/rashomon, else ~/.local/state/rashomon. RASHOMON_STORE_CAP_BYTES (default 512 MiB) is a soft limit: at session start and end, the least recently written runs are evicted, each leaving a gap record so the deletion is visible. The session whose start or end triggered the pass, and any run written in the last hour, are never evicted, so the store can run past the cap for a while. Another session left open but idle for more than an hour is not protected. detach removes the hooks, not the records. To remove everything, delete that directory.

Redaction, and what it does not hide

report --redact digests hostnames under this install's key. It is for sharing a report, and its limits are real: the digest is 32 bits, so two hosts can collide; the last label is kept in clear on purpose, so .internal and .com survive (x.s3.amazonaws.com keeps only .com), and an IPv4 address keeps its last octet; and anyone holding this install's store can recompute every digest. The text report repeats all of this in its own header; --json --redact marks only "redacted": true.

forget --host <h> removes every call that named that host from this store, and leaves a gap record saying something was removed. Pass the host exactly as the report prints it (lower-case, no scheme or port): in this release any other form matches nothing and reports forgot 0 records. That record is keyed by a digest of the host, not its name.

What it cannot see

Network destinations are not observed in this alpha. The text report says so in one line:

destinations: not observed in this alpha

It is printed for every session, including a completely healthy one, because it is a property of the instrument and not of the session. A reader told only what was observed will read the rest as an absence of traffic rather than an absence of observation. --json states the same fact in its own form: the destinations object carries "observed": false and the reason no_proxy_store, so its empty host lists read as not watched, not as none. --chain lists the hosts each call named; it does not say whether the call reached them.

The two rules it will not break

Never print "nothing happened" where the truth is "not watching." Every count that could not be measured renders as unknown, never as zero. A run that could not measure itself does not get to report a number. One known exception in this release: executed differently from declared prints 0 even when no declared and executed pair could be compared.

Never print "not watching" where the truth is "nothing happened." Every degradation line has a healthy twin, so a clean run and a blind run do not read the same.

One consequence worth expecting: coverage reads verified only when the probe can find the recorder's entry, either in the settings file Claude Code itself resolves (a watch install) or as the enabled plugin. A session started with claude --settings <other file> records perfectly well and still reports hook_entry_absent: the probe under-claims rather than confirming an entry it cannot read.

The reasoning behind all of this, and the acceptance criteria that hold it, is in docs/design-notes.md.

Inside a sandbox

If Claude Code runs inside a sandbox that limits where it can write, the sandbox must allow rashomon's store directory, or nothing is recorded and nothing says so. With nono, create the directory first, then grant it: mkdir -p ~/.local/state/rashomon, and add --allow ~/.local/state/rashomon to nono run. Running Claude Code under nono has not been tested end to end in this release.

Commands

Command What it does
rashomon watch Install the recorders, the liveness probe and the end-of-turn recap (eight entries). The only command that installs anything.
rashomon report [--session S] [--json] [--redact] [--chain] [--timeline] Render every recorded session, or one named with --session. --chain adds the causal view: which prompt produced which calls. --timeline lists every call, main agent and subagents, in the order they were recorded, keeps failed calls apart from calls that never ran, and says whether a success of the same command, or of the same program for single-purpose programs, was recorded after each failure, and reads "not checked" where a success may exist but cannot be placed or matched, or where the failure itself has no declaration or recorded position; for wrappers and multi-command programs such as git, go, make, npm, python and sudo, and for calls with no program such as Read or Edit, only the same command is looked for, and a fix made with a different command, or a corrected Edit, is not detected
rashomon status Say what is installed here. Reads only; creates nothing
rashomon spend [--days N] [--json] Estimate what the last N days of Claude Code usage would cost at API list prices, from Claude Code's own transcripts. Needs no watch; writes nothing. See rashomon spend
rashomon pause / rashomon resume Stop and restart recording on this machine, leaving a record of the change
rashomon detach Remove the recorders, leaving every other entry's value as found
rashomon forget --since T | --before T | --host H Erase records, leaving a gap record saying so
rashomon version Print the version

rashomon --help lists the commands above and their main flags. It does not list detach --force, the override a refused detach names in its error message. The commands in the table do not take --help in this release: each refuses it, like any argument it does not know, with unknown argument and exit status 1, and does nothing else, unless it comes where a flag expects its value.

detach has two recovery forms for when things are gone: detach --install <id> (printed by watch at install time) removes one install's entries without opening a store, and detach --all removes every entry carrying a rashomon marker, for when the store is gone and the id with it. Both refuse, naming the field, if an entry of ours had its matcher, hook count, hook type or timeout edited by hand.

rashomon spend

rashomon spend [--days N] [--json] estimates what the last N days (default 30, at most 36500) of Claude Code usage would cost at API list prices. It needs no watch, and it writes nothing.

Not a bill. The rates are a dated snapshot compiled into the binary, and every rendering names the date. A Pro or Max plan is not billed per token, so the figure is what the same usage would cost on the API, not what you were charged.

What it reads. Claude Code's own transcripts, under $CLAUDE_CONFIG_DIR/projects or ~/.claude/projects, subagent transcripts included. The totals come from each API response's usage fields only: its id, model, stop reason, a refusal's category, token counts and speed, each retry attempt's counts, type and model, and the line's timestamp, session, isSidechain and request id. Claude Code writes one response on several lines, so a response is counted once by its id.

What it shows. Spend by agent, model, token kind and session. Cache re-written after a gap longer than its TTL: the part of a write that re-writes what the previous response had cached and this one did not read back, priced over a cache read. It is a heuristic, labelled as one: it skips a model switch and any request smaller than the previous cache (after compaction, say), since it wrote new content, so it misses some true expiries, and can still count new content in a request that grew past the previous cache. Refusals, counted by category and model. Retry attempts: every attempt before the last (the last produced the message and is the top-level usage) is shown in tokens with the cost unknown and left out of the total, and the header says so. When the last attempt is a fallback model's (fallback_message), the response is priced at that model, and counted as one a fallback model served unless it ended in a refusal. And the spend in turns with a failed call the summary never mentioned.

Refusals. A refusal's stop_details.category is read as a closed word: cyber, bio, frontier_llm, reasoning_extraction, general_harms, uncategorized (null) or other, and refusals are counted by category and model. A refusal partway through its output (output_tokens above 0) is billed at normal rates, and priced like any response. Whether one before any output (output_tokens 0) was billed depends on its category, so its tokens are shown with the cost unknown and left out of the total, and the header says how many there are. After a refusal, Claude Code also writes one line whose usage is all zeros (model <synthetic>) with the same requestId as the response; that line is folded into the response. One with no such response is counted as a pre-output refusal written without usage.

The silent-failure line. A turn counts when one of its recorded calls failed and its final message mentions no failure, whether or not a later call succeeded. A failed call whose declaration was lost or carried no prompt_id has no prompt, so it is placed in no turn: the line counts it as a failed call that could not be checked, and with one, "none found" holds only for the turns that could be. The line covers only the transcripts a rashomon record names, and the transcript of a session rashomon watched with no tool call made, when it was read whole, no response read from it asked for a tool, and every response in the window came while rashomon was watching. A session with no tool call that is still open has no end record yet, and reads not covered until it ends. Every other transcript is counted and its spend priced and marked not covered, never folded in as zero, and the sessions whose rows hold that spend are named. A session's row reads recorded only when its own transcripts were recorded and it holds none of that spend. A response that appears in several transcripts (a resumed conversation, or a copy made by /branch or --fork-session) is not covered when any of them is not, and its cost is counted once. In the per-session rows it belongs to the session whose transcript starts first; a copy that keeps the original's timestamps starts at the same moment, and then both sessions' rows hold it; the output then says how many such responses there are (shared_responses in --json), since the rows can add up to more than the total. A turn's spend is the responses tied to its prompt. In the main transcript, and in each subagent transcript under it, a response belongs to the prompt of the user line before it: Claude Code writes that prompt's id (promptId) on the line. The figure is a floor. A response after a user line with no prompt id belongs to no turn, unless that line is an injected meta line or, in the main transcript, a tool result. A subagent's response written into the main transcript (isSidechain) counts toward the prompt before it, and a subagent's user line there does not end the tie. A line that cannot be decoded ends it (any in a subagent transcript; in the main transcript, unless it is a sidechain line), so a response after it is not counted. A failed turn whose final message cannot be tied to its prompt is counted as not checked, never as clean.

Message content. To take that verdict, the line reads message content, in memory. For each recorded turn with a failed call, it decodes each block's type, and a text block's text, of every assistant line tied to the turn, and keeps only the last line's text. The text is never written or output. To tell a prompt from a tool result on a user line with no prompt id, it decodes the line's content block types only, never their text. From every line of a subagent transcript, and from a main transcript's lines up to its first dated one, it decodes the fields listed under What it reads and the line's type, isMeta and promptId, never its content.

Savings. One saving is listed, with its figure: the spend in turns with a failed call the summary never mentioned. The cache re-write figure is a heuristic, so it is shown and not offered as a saving. Pre-output refusals and retry attempts are tokens with the cost unknown, so they carry no saving.

Counted, said, never priced. It says how many of each of these it found: lines it cannot count (malformed or implausible usage, a line with no message id, a future-dated response), and folders and files it cannot read to the end. Fast-mode responses are counted and priced at standard rates. Out of scope: fast mode's premium, Batch and partner (Bedrock, Vertex) pricing, long-context premiums, and web-search fees.

Editor integration

This repository installs nothing in your sessions. There is no committed .claude/settings.json; opening this checkout runs nothing. The recorder runs only once you run watch or enable the plugin.

Slash commands ship as assets you copy from a clone of this repository (go install leaves none), deliberately not under .claude/ here. A skill in this repository would load only for someone who has this repository open, which is the one place it is least useful:

mkdir -p "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/skills" && cp -r skills/rashomon* "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/skills/"   # every project
mkdir -p /path/to/project/.claude/skills && cp -r skills/rashomon* /path/to/project/.claude/skills/                            # one project

scripts/claude-rashomon starts a recorded session in one word once it is on your PATH. It runs watch, which installs user-wide hooks that keep recording every later session, in every project, until detach. scripts/rashomon-sessionstart.sh is an optional SessionStart hook that puts one line of recorder state into a session's context; its header shows the entry to add. It reads only a settings install, so with the plugin it wrongly reports the session as not recorded.

Scope and threat model

Claude Code only. The hooks, the store and the report describe Claude Code sessions. Other tools that read ~/.claude/settings.json can trigger the recorder anyway: Cursor by default, and Codex after /import. This release does not support them, and what gets recorded for them can be wrong. If you use Cursor with rashomon installed, turn off Include Third-Party Plugins, Skills, and Other Configs (Cursor Settings, Agents, Third-Party Imports).

rashomon installs no certificate and decrypts nothing.

No account. No telemetry. No phone-home. No model in any path.

What it defends against, stated plainly: this release defends against the agent's narration: a summary that omits, compresses or misremembers. It does not defend against an agent that tampers. The store, the digest key and forget are all reachable by any shell command the agent makes. Treat the record as an honest second account, not as tamper-proof evidence.

A declaration is not an execution. PreToolUse fires before the permission prompt resolves, so the store holds records for calls you went on to deny. That is deliberate (intent is worth keeping whether or not it was granted). Each declaration carries the session's permission_mode (for example default or acceptEdits). That is the mode, not the answer you gave the prompt, so it cannot by itself separate what ran from what you refused.

Building from source

go build ./cmd/rashomon
go test ./...

Status

Alpha: v1.1.0. No signed binaries, no Homebrew formula. macOS and Linux are the supported targets.

Windows is not usable yet. The release publishes Windows binaries, but no install path has been shown to work there. watch quotes the installed command line for a POSIX shell, and scripts/claude-rashomon.ps1 runs the same watch. CI cross-builds and vets the Windows code; no test runs on Windows.

Contributing

Issues and pull requests are welcome. main is protected: changes land through a pull request with CI green. Security issues should go through private vulnerability reporting rather than a public issue; see SECURITY.md.

License

Apache 2.0. See LICENSE.


Built by Altrace.

来源:Hacker News · AI · github.com