Repository navigation
Conversation
…-network/harness-sessions readClaudeTranscript and parseClaudeTranscript project the normalized session of the shared reader: one assistant message and one usage count per API response. The old line parser counted a response's usage again when its first block was an empty (redacted) thinking block, up to 2x on real transcripts. The OpenCode reader reads a private copy of the store and no longer refuses sessions with an errored or unfinished tool part; its API takes the store path. The supervision-tree Claude reader reads tool traffic, structured results and notifications from the same session. Intake Claude metrics take token totals from the shared fold.
fromPiSession reads a native Pi session (~/.pi/agent/sessions) through the shared reader: tool-call actions with their results, the last answer, the ending, tokens and the cost Pi records per call, and the served provider/model. The graph IR it used to parse is fromPiGraphSession (source pi-graph). Before this, nothing in agent-eval could read a real Pi session.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
agent-eval kept its own Claude Code and OpenCode session parsers, separate from every other consumer. On 728 stored Claude Code transcripts (GTR, 2026-10-04) its rollout reader over-counted usage on 346: a response written as several records is counted again when its first block is an empty (redacted) thinking block, up to 2x tokens. Re-simulating that rule reproduces its numbers exactly on 619/619 sessions. The intake summed usage on every content-block record. The OpenCode reader threw on 10 of 26 real sessions (any errored or unfinished tool part).
Change
rollout/readers/claude-jsonl.ts:readClaudeTranscript/parseClaudeTranscriptproject@tangle-network/harness-sessionssessions (transcriptFromSession). One assistant message per API response (its tool calls together, then their results); usage once per response; a usage field is null when any answered call lacked it; the model is the last served model (never<synthetic>).parseClaudeEntries/transcriptFromEntriesare deleted.rollout/readers/opencode-sqlite.ts:findOpencodeSessionsByDirectory(directory, db),readOpencodeSession,readOpencodeSessionMessages(sessionId, db)over a private copy of the store.openOpencodeDbandOpencodeSessionRoware removed (no consumers found).supervisor-run/claude-code-reader.ts: tool traffic, structured results (toolUseResultasresult.details) and task notifications come from the same session.contract/intake/code-agent-session.ts: Claude Code token totals come from the shared fold.node:sqlitelazily).Proof
tsc --noEmit,biome check src, and the affected suites (rollout readers, supervisor-run incl. the 52-agent Claude Code fixture, contract intake): 217 tests pass (incl. a native Pi intake test on a real Pi 0.85.1 session) on beelink1 against the packed package.contract/intake:pinow means a native Pi session, read through the shared reader (tool-call actions, last answer, ending, tokens, Pi's per-call cost, served provider/model); the graph IR it used to parse isfromPiGraphSession(sourcepi-graph). Before this, nothing in agent-eval could read a real Pi session.Remaining in this repo
The intake projections for Codex (rollout and
exec --jsonstream), Kimi, and OpenCode'srun --format jsonstream still parse records themselves; the Claude projection still counts actions from records. Kimi's intake metrics already agree with the shared reader on 3,000/3,000 stored sessions.Blocked
@tangle-network/harness-sessions@0.1.0awaits its first npm publish. After it publishes: add the dependency (^0.1.0), update the lockfile, frozen-install gate, merge, release.