fix(ai): bound cumulative tool results per request - #464
Conversation
每次模型请求前检查总预算并裁剪工具结果副本,保留原记录、调用配对和计费,避免并行或跨轮结果累积超窗。 Fit cumulative tool results before each model request while preserving the transcript, call pairing and usage. Verified with 121 related tests, 40 stress scenarios, Pi HTTP failure injection and independent reviews.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 53e664d910
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| : message | ||
| ), | ||
| }, | ||
| maxTokens: Math.max(1, Math.min(maxTokens, contextWindow - tokens - 16)), |
There was a problem hiding this comment.
Preserve the Responses API's 16-token minimum
When an openai-responses request has only 16–31 tokens left after tool-result trimming—reachable when protected user or history text occupies most of the window—this expression sends maxTokens between 1 and 15 because it subtracts 16 and then floors at 1. The preceding comment documents that this provider requires at least 16 output tokens, so the request is rejected even though a 16-token response would fit. Clamp the output budget to at least 16 while respecting the actual remaining window, and cover this boundary through the Responses adapter.
Useful? React with 👍 / 👎.
|
目前和最新代码有一些冲突了,麻烦解决一下 |
Merge main b233725; preserve tool-result budgeting, output-truncation handling, and both regression suites. 合入最新 main,保留工具结果预算、输出截断处理及双方回归测试。
|
已合入最新 main(b233725d),冲突已解决:保留工具结果预算保护、上游输出截断处理及双方测试。84 项相关测试、40 组压力测试(120 次请求边界)及 Node/Desktop 类型、lint、格式检查均通过。提交:951d391e。 Merged latest main (b233725) and resolved the conflicts, preserving tool-result budgeting, upstream truncation handling, and both test sets. Passed 84 related tests, 40 stress scenarios (120 request boundaries), Node/Desktop type checks, lint, and formatting. Commit: 951d391. |
|
已同步最新 main(a67e38b8),包含您在 45cfe87 补齐的时间参数英文说明。此前失败的本地化检查已通过,本轮 CI:2385 项测试全部通过,类型检查与 Web WASM 构建通过。更新提交:f74e71d2。 Synced latest main (a67e38b), including your metadata fix in 45cfe87. The previously failing localization check now passes; CI passed all 2,385 tests, type checks, and the Web WASM build. Updated commit: f74e71d. |
原因 / Why
单个工具的截断预算无法限制并行或多轮结果的总量,下一次模型请求仍可能超窗失败。
Per-tool limits do not bound accumulated results, so parallel or successive tool calls can overflow the next request.
改动 / Changes
共享请求入口计入系统提示、历史和工具定义;优先裁剪较早的工具文本副本,保留配对、原始记录及计费。输出上限按剩余空间计算,避免中文历史被压到 1 token。
Budget each request and trim older tool text in its request view. Preserve call pairs, original results and usage; invalidate stale prefix estimates only in the edited view. Cap output by remaining space without starving CJK continuations.
验证 / Verification
git diff --checkpassed.pnpm test -- packages/node-runtime/src/ai/agent/__tests__/core.test.ts.边界 / Limits: uses conservative cl100k/character estimates, not exact tokenizers for every provider. Trimming reduces older evidence in requests; raw tool results remain intact. Tests use synthetic data/local HTTP, not paid model calls.