Sitelet https://github.com/braintrustdata/autoevals/pull/232
Skip to content

feat: add voice task success scorer - #232

Draft
kev (kevchoi) wants to merge 1 commit into
mainfrom
kevchoi/20261002-voice-task-success-sdk
Draft

kev (kevchoi) wants to merge 1 commit into
mainfrom
kevchoi/20261002-voice-task-success-sdk

Conversation

@kevchoi

@kevchoi kev (kevchoi) commented Oct 2, 2026 •

Copy link
Copy Markdown

AI-created or modified and not human-reviewed in its current form; treat this artifact as provisional. Updated 2026-10-02.

Summary

Adds VoiceTaskSuccess, an LLM judge for whether a voice agent completed the caller's request, judged from the whole call (instructions, tool calls, tool results).

  • templates/voice_task_success.yaml: A/B/C rubric over {{thread_with_system}}.
  • JS and Python VoiceTaskSuccess: picks the conversation from a thread_with_system argument, else trace.getThread() / trace.get_thread(). With no conversation it returns a null score without calling the model. Otherwise it passes the messages to the template as a JSON string, so tool calls and results are included.
  • js/manifest.ts entry, so it's available as a built-in scorer.
  • SCORERS.md, README.md, AGENTS.md.

Testing

  • Build, tsc --noEmit (one pre-existing error in js/render-messages.test.ts), vitest, pre-commit. Tests that need an OpenAI key fail the same way on main.
  • Ran JS and Python (sync and async) against a live model for each source: thread_with_system, trace, no conversation (null), empty trace (null), and .partial.
  • JS JSON.stringify and Python json.dumps produce identical text for the same conversation. The judge cites tool names and arguments (a transfer of $1000 to Bob when the caller asked for $10 to Alice scores 0).

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (HEAD-1790984589)

Score Average Improvements Regressions
NumericDiff 77.6% (-2pp) 4 🟢 7 🔴
Time_to_first_token 7.33tok (-0.53tok) 142 🟢 77 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 512.4tok (-2.98tok) 26 🟢 18 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 461.71tok (-23.34tok) 111 🟢 101 🔴
Completion_reasoning_tokens 341.24tok (-24.44tok) 88 🟢 74 🔴
Completion_accepted_prediction_tokens 0tok (+0tok) - -
Completion_rejected_prediction_tokens 0tok (+0tok) - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 974.11tok (-26.32tok) 113 🟢 100 🔴
Estimated_cost 0$ (0$) 70 🟢 59 🔴
Duration 7.33s (-0.53s) 142 🟢 77 🔴
Llm_duration 8.15s (-0.6s) 140 🟢 79 🔴

@kevchoi
kev (kevchoi) force-pushed the kevchoi/20261002-voice-task-success-sdk branch 3 times, most recently from bca422d to 9a9e057 Compare October 2, 2026 22:37
@kevchoi
kev (kevchoi) force-pushed the kevchoi/20261002-voice-task-success-sdk branch from 9a9e057 to e2e2d83 Compare October 2, 2026 22:55
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevchoi
kev (kevchoi) force-pushed the kevchoi/20261002-voice-task-success-sdk branch from e2e2d83 to b0598fe Compare October 2, 2026 23:42

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant