Sitelet https://github.com/lidge-jun/opencodex/pull/6741
Skip to content

fix(claude): carry a provider's own encrypted reasoning through the Claude route - #6741

Draft
rhomat27 wants to merge 1 commit into
lidge-jun:devfrom
rhomat27:fix/claude-native-reasoning-continuity
Draft

rhomat27 wants to merge 1 commit into
lidge-jun:devfrom
rhomat27:fix/claude-native-reasoning-continuity

Conversation

@rhomat27

@rhomat27 rhomat27 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes the reasoning loss reported in #6736. On the translated Claude Messages → Responses route, a Responses provider's own encrypted_content was dropped on the way out (outbound.ts decodes only ocxr1: envelopes). Nothing could carry it back either: the body is store: false and Claude Code replays only the thinking block. So the routed model lost its own reasoning every turn. With Muse Spark that shows up as long reading phases before the first edit, with the model restating that it is ready many times.

  • src/responses/reasoning-envelope.ts: new envelope field nat holds the provider's blob verbatim, plus the client-facing model it was minted for and the provider's item id. It follows the same pattern as krc for Kiro. A malformed nat is ignored.
  • src/claude/outbound.ts (streaming and JSON): a reasoning item with a native blob is signed into its thinking block's envelope. It now gets a thinking block even when no summary arrived; before, most Muse items produced no block at all.
  • src/claude/inbound.ts: a replayed thinking block whose envelope has nat becomes a reasoning item with that encrypted_content and id, but only when the request names the same model. Any other model gets the visible text alone, so /model mid-conversation can't hand a blob to a provider that can't decrypt it. Bodies that reason now send include: ["reasoning.encrypted_content"], the same field Codex sends to every Responses provider.
  • Anthropic signatures (sig) and redacted blocks (red) behave as before. Nothing changes for native Anthropic passthrough.
  • structure/clients/claude-desktop.md documents the rule. tests/claude-integration/claude-native-reasoning-continuity.test.ts is registered in the test layout on an existing line.

This covers the Responses-route fix in #6736. The issue also has a comment about Meta's Anthropic-format endpoint through the anthropic adapter (src/protocols/opaque-state.ts); that one is a separate design question and is not touched here.

Closes #6736

Verification

Run on macOS 26.6.2 (arm64), Bun 1.4.0, branch rebased onto dev at bf9ecf3.

  • bun run typecheck: pass.
  • bun test tests/claude-integration/claude-native-reasoning-continuity.test.ts: 10 pass. With src/ reverted to dev, 8 of the 10 fail. The 2 that pass either way check unchanged behaviour (an item with no native blob, a malformed nat).
  • Neighbouring suites: claude-outbound, claude-inbound, claude-inbound-tool-reference, claude-inbound-token-footer, responses/reasoning-envelope, providers/kiro/kiro-reasoning-roundtrip, adapters/bridge, adapters/anthropic/anthropic-opaque-strip. Together with the new file: 338 pass, 0 fail.
  • bun run structure:check: pass. bun run privacy:scan: pass.
  • bun run test:changed (before the rebase, dev at 18a8cda): 33,105 tests, 32,798 pass, 201 fail, 3 errors. Every failure sits in 26 files outside the changed modules (server auth and admission, lab, Kiro catalog, service).
    • Running those 26 files on an unmodified dev worktree gives the same failing test names: 656 pass, 183 fail in isolation on both trees.
    • This workstation runs a live opencodex and inherits Claude Code's HTTPS_PROXY and NODE_EXTRA_CA_CERTS. With those cleared, the same files drop to 56 failures, again identical on dev. I'm leaving the full-suite verdict to CI.

Live check: the same diff applied to an installed v2.80.0, route meta-muse/muse-spark-1.3-contributor (Coding Plan).

  • Recall test. The model is told a random 8-character code and asked to keep it only in its reasoning. On turn 2 the code is replaced with [removed] and the model is asked for it.

    Turn 2 Recalled
    Same model, fix applied 3 of 5
    Control: same prompt, same provider, turn 2 sent as the [1m] alias, so the guard withholds the blob 0 of 5

    Without the fix this route recalled 0 of 3 on the earlier test.

  • A real Claude Code subagent on the Muse route, given a small test-writing task:

    • all 7 of its thinking blocks carried the native blob
    • first file edit at 15 s
    • no proxy errors
    • all 27 Muse requests in that window reported reasoning tokens. On our 2.76 install's log, only about 10–18% of Claude-route Muse requests did.

Checklist

  • Scope stays focused and avoids unrelated cleanup.
  • Docs or release notes were updated when needed (structure/clients/claude-desktop.md).
  • Security-sensitive changes were reviewed for secrets, auth, and unsafe defaults. The blob is the provider's own opaque response data, already returned to this proxy. It goes back only to the model that produced it, and no credential or account data enters the envelope.

🤖 Generated with Claude Code

Review readiness checklist

This PR stays in draft until every box below is ticked. Tick all four boxes once the requirements are met:

  • Required local validation passed; commands, results, and any full-suite exception are documented.

  • I pushed my PR to a recent dev commit (at most 10 behind; a maintainer may still ask for the exact tip before merge).

  • I resolved all correct Codex and CodeRabbit findings.

  • My PR is ready for review.

Summary by CodeRabbit

  • New Features

    • Preserved encrypted reasoning across Claude and Responses conversions, including streaming and JSON responses. Reasoning is replayed only when the requested model matches; otherwise, visible thinking text is retained when available.
    • Requests with thinking or output-effort settings now include encrypted reasoning in the response.
  • Tests

    • Added coverage for reasoning continuity, model matching, empty thinking blocks, and malformed reasoning data.

…laude route

On the translated Claude Messages route a Responses provider's native
encrypted_content was discarded on the way out, and nothing could carry it
back: the body is store:false and Claude Code replays only the thinking
block. The routed model lost its own reasoning every turn (lidge-jun#6736).

Keep the native blob in the thinking block's ocxr1 envelope as `nat`, with
the client-facing model and the provider's item id, and emit that block even
when no summary arrived. Inbound restores it as encrypted_content and id
only for a request naming the same model; any other model gets the visible
text alone. Translated bodies that reason now request
include: ["reasoning.encrypted_content"], as Codex does.

Refs lidge-jun#6736

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the bug Something isn't working label Oct 8, 2026
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: lidge-jun/opencodex/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 74b1d7fa-2600-4571-ae34-dfdaa0dfa0ec
📥 Commits

Reviewing files that changed from the base of the PR and between bf9ecf3 and 62401f2.

📒 Files selected for processing (7)
  • scripts/test-layout/layout.json
  • src/claude/inbound.ts
  • src/claude/outbound.ts
  • src/responses/reasoning-envelope.ts
  • structure/clients/claude-desktop.md
  • tests/claude-integration/claude-native-reasoning-continuity.test.ts
  • tests/fixtures/test-layout-expected.json

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The Claude translation now preserves provider-native encrypted reasoning in Claude thinking signatures and restores it in Responses requests for the same model. It also requests encrypted reasoning when thinking or output effort is configured. New integration tests cover conversion, replay, and malformed envelopes.

Changes

Claude reasoning continuity

Layer / File(s) Summary
Native reasoning envelope
src/responses/reasoning-envelope.ts
The envelope supports opaque native reasoning content with a model and optional item ID. Decoding accepts only nonempty encrypted content and model strings.
Responses-to-Claude preservation
src/claude/outbound.ts, tests/claude-integration/claude-native-reasoning-continuity.test.ts
Streaming and JSON translations preserve qualifying native encrypted content in thinking signatures, including when summary text is absent. Tests cover native and summary-only reasoning.
Claude-to-Responses replay
src/claude/inbound.ts, tests/claude-integration/claude-native-reasoning-continuity.test.ts, structure/clients/claude-desktop.md, scripts/test-layout/layout.json, tests/fixtures/test-layout-expected.json
Inbound translation restores the encrypted content and item ID only when the requested model matches. Configured thinking or output effort requests reasoning.encrypted_content. Tests cover replay, request configuration, and invalid envelopes; documentation and test-layout mappings were updated.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix · Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to 62401

No concrete issue remains that would prevent merging after normal checks.

🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning Issue #6736 requires native reasoning to return only to the same provider and model. The change stores only enc, model, and optional id in ReasoningEnvelope.nat (`src/responses/reasoning-envel… Store a stable provider identity with nat when the provider creates the reasoning item. Restore encrypted_content and id only when both that provider identity and the requested model match the current route. Otherwise retain only visi…
Docstring Coverage ⚠️ Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 14 functions across 4 files. (3 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: preserving a provider's encrypted reasoning through the Claude route. It matches the implementation and PR objective.
Out of Scope Changes check ✅ Passed The changed source files implement the linked reasoning-continuity behavior. The new integration tests cover outbound preservation, inbound replay, model mismatch, empty thinking, request inclusion, m…
Full details: Linked Issues check

Explanation

Issue #6736 requires native reasoning to return only to the same provider and model. The change stores only enc, model, and optional id in ReasoningEnvelope.nat (src/responses/reasoning-envelope.ts, NativeReasoning and decodeNativeReasoning). The inbound path receives and compares requestedModel, but it has no provider identity to store or compare (src/claude/inbound.ts, assistantMessageToItems). A request for the same client-facing model that routes to a different Responses provider can therefore receive the original provider's opaque blob and item ID. The outbound changes correctly preserve unprefixed blobs and emit empty thinking blocks (src/claude/outbound.ts, responsesSseToAnthropicSse and responsesJsonToAnthropicMessage), and the tests cover the model mismatch, but they do not establish provider matching.

Resolution

Store a stable provider identity with nat when the provider creates the reasoning item. Restore encrypted_content and id only when both that provider identity and the requested model match the current route. Otherwise retain only visible thinking text, or omit the reasoning item when that text is empty. Add a test for the same model routed to a different provider.

Full details: Docstring Coverage

Explanation

Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 14 functions across 4 files. (3 skipped: 3 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

⏳ DRAFT

  • review readiness checklist open (0/4 boxes ticked).

What to do

  • Tick all four boxes in the PR description once you're done (currently 0/4).

Review readiness checklist

  • ⬜ Required local validation passed; commands, results, and any full-suite exception are documented.
  • ⬜ I pushed my PR to a recent dev commit (at most 10 behind; a maintainer may still ask for the exact tip before merge).
  • ⬜ I resolved all correct Codex and CodeRabbit findings.
  • ⬜ My PR is ready for review.

0/4 boxes ticked.

This PR stays in draft until every box above is ticked.

@github-actions
github-actions Bot marked this pull request as draft October 8, 2026 04:03
@rhomat27

rhomat27 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Results from a second machine, plus prompt caching, which the PR description didn't cover.

The same diff runs on v2.80.0 on two machines, routing Claude Code subagents to meta-muse/muse-spark-1.3-contributor through the first-party intercept: the laptop from the description, and a Mac mini that serves a second host over the LAN with stabilizePromptCache on. On the Mini it runs alongside a cherry-pick of #6714.

  • Recall test on the Mini (same test as in the description): 5 of 8. Before this change, that route recalled 0 of 3.
  • Reasoning: all 37 Muse requests on the Mini in the first minutes after the restart reported reasoning tokens. On the same machine before the change, about 10–18% of Claude-route Muse requests did.
  • Prompt caching is unaffected. Replaying the blob and sending include don't break prefix caching.
    • Laptop: Muse follow-up turns read 90% of their input from cache, across 39 turns.
    • Mini: in a new conversation after the restart, turns 3 and 4 read 99% and 98% from cache, and turn 5 read everything already sent (56.7k of 77.1k input), so only its new content was uncached.
  • Model switch: sending turn 2 under a different model name keeps the blob back with no error (laptop: 0 of 5 recalled, as intended). On the Mini, opencodex resolves the [1m] alias to the same model before translation, so the blob is forwarded there.

🤖 Generated with Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant