Sitelet https://github.com/tangle-network/agent-eval/pull/938
Skip to content

feat(readers): read Claude Code and OpenCode sessions through @tangle-network/harness-sessions - #938

Draft
drewstone wants to merge 2 commits into
mainfrom
feat/harness-sessions-readers-20261004
Draft

drewstone wants to merge 2 commits into
mainfrom
feat/harness-sessions-readers-20261004

Conversation

@drewstone

@drewstone drewstone commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Problem

agent-eval kept its own Claude Code and OpenCode session parsers, separate from every other consumer. On 728 stored Claude Code transcripts (GTR, 2026-10-04) its rollout reader over-counted usage on 346: a response written as several records is counted again when its first block is an empty (redacted) thinking block, up to 2x tokens. Re-simulating that rule reproduces its numbers exactly on 619/619 sessions. The intake summed usage on every content-block record. The OpenCode reader threw on 10 of 26 real sessions (any errored or unfinished tool part).

Change

  • rollout/readers/claude-jsonl.ts: readClaudeTranscript / parseClaudeTranscript project @tangle-network/harness-sessions sessions (transcriptFromSession). One assistant message per API response (its tool calls together, then their results); usage once per response; a usage field is null when any answered call lacked it; the model is the last served model (never <synthetic>). parseClaudeEntries / transcriptFromEntries are deleted.
  • rollout/readers/opencode-sqlite.ts: findOpencodeSessionsByDirectory(directory, db), readOpencodeSession, readOpencodeSessionMessages(sessionId, db) over a private copy of the store. openOpencodeDb and OpencodeSessionRow are removed (no consumers found).
  • supervisor-run/claude-code-reader.ts: tool traffic, structured results (toolUseResult as result.details) and task notifications come from the same session.
  • contract/intake/code-agent-session.ts: Claude Code token totals come from the shared fold.
  • Tests run on sessions the real CLIs wrote against the trace-proof scripted model; the module-mock lazy-load test is deleted (the package loads node:sqlite lazily).

Proof

tsc --noEmit, biome check src, and the affected suites (rollout readers, supervisor-run incl. the 52-agent Claude Code fixture, contract intake): 217 tests pass (incl. a native Pi intake test on a real Pi 0.85.1 session) on beelink1 against the packed package.

  • contract/intake: pi now means a native Pi session, read through the shared reader (tool-call actions, last answer, ending, tokens, Pi's per-call cost, served provider/model); the graph IR it used to parse is fromPiGraphSession (source pi-graph). Before this, nothing in agent-eval could read a real Pi session.

Remaining in this repo

The intake projections for Codex (rollout and exec --json stream), Kimi, and OpenCode's run --format json stream still parse records themselves; the Claude projection still counts actions from records. Kimi's intake metrics already agree with the shared reader on 3,000/3,000 stored sessions.

Blocked

@tangle-network/harness-sessions@0.1.0 awaits its first npm publish. After it publishes: add the dependency (^0.1.0), update the lockfile, frozen-install gate, merge, release.

…-network/harness-sessions

readClaudeTranscript and parseClaudeTranscript project the normalized
session of the shared reader: one assistant message and one usage count
per API response. The old line parser counted a response's usage again
when its first block was an empty (redacted) thinking block, up to 2x on
real transcripts. The OpenCode reader reads a private copy of the store
and no longer refuses sessions with an errored or unfinished tool part;
its API takes the store path. The supervision-tree Claude reader reads
tool traffic, structured results and notifications from the same
session. Intake Claude metrics take token totals from the shared fold.
fromPiSession reads a native Pi session (~/.pi/agent/sessions) through
the shared reader: tool-call actions with their results, the last
answer, the ending, tokens and the cost Pi records per call, and the
served provider/model. The graph IR it used to parse is fromPiGraphSession
(source pi-graph). Before this, nothing in agent-eval could read a real
Pi session.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant