Sitelet https://github.com/braintrustdata/autoevals/pull/228
Skip to content

feat: add speech clarity and turn-taking voice scorers - #228

Closed
kev (kevchoi) wants to merge 3 commits into
mainfrom
kevchoi/20260928-autoevals-voice-scorers
Closed

kev (kevchoi) wants to merge 3 commits into
mainfrom
kevchoi/20260928-autoevals-voice-scorers

Conversation

@kevchoi

@kevchoi kev (kevchoi) commented Sep 30, 2026 •

Copy link
Copy Markdown

AI-created or modified and not human-reviewed in its current form; treat this artifact as provisional. Updated 2026-10-02.

Exploratory draft, not ready for review.

Context

A first pass at voice scorers. Both are LLM-as-a-judge scorers that listen to a recording of the whole conversation and judge only the agent.

Description

  • Adds SpeechClarity and TurnTaking in Python and JS, with templates in templates/speech_clarity.yaml and templates/turn_taking.yaml, listed under "LLM-as-a-Judge" in js/manifest.ts.
  • Adds an optional audio: template field, a dotted path such as input.audio. It points to {data: bytes, content_type}. Missing audio skips the score (score=None), and a non-audio/* content type raises an error.
  • Sends the audio unchanged as a chat file part with a base64 data URL. Only Gemini models read it this way. The default model is gemini-3.8-flash.
  • An earlier version converted OGG to 128 kbps MP3 for input_audio. On damaged test clips the MP3 path scored clean calls as flawed and missed most turn-taking failures, while sending OGG unchanged did at least as well as WAV. So the conversion and its dependencies were removed.

Testing

  • Mocked request-shape tests: js/llm.test.ts and py/autoevals/test_llm.py -k speech_clarity.
  • Live runs against Gemini on clean and deliberately damaged sample calls.

🤖 Generated with Claude Code

kev (kevchoi) and others added 2 commits September 30, 2026 13:11
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Braintrust eval report

Autoevals (HEAD-1790808392)

Score Average Improvements Regressions
NumericDiff 78.7% (+1pp) 9 🟢 9 🔴
Time_to_first_token 8.54tok (-1.75tok) 180 🟢 39 🔴
Llm_calls 1.55 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 515.75tok (-0.75tok) 26 🟢 24 🔴
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 465.66tok (+0.05tok) 110 🟢 102 🔴
Completion_reasoning_tokens 349.67tok (-0.58tok) 93 🟢 86 🔴
Completion_accepted_prediction_tokens 0tok (+0tok) - -
Completion_rejected_prediction_tokens 0tok (+0tok) - -
Completion_audio_tokens 0tok (+0tok) - -
Total_tokens 981.41tok (-0.69tok) 111 🟢 101 🔴
Estimated_cost 0$ (+0$) 68 🟢 68 🔴
Duration 8.55s (-1.75s) 180 🟢 39 🔴
Llm_duration 9.31s (-3.71s) 190 🟢 29 🔴

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevchoi

Copy link
Copy Markdown
Author

AI-created or modified and not human-reviewed in its current form; treat this artifact as provisional. Updated 2026-10-02.

Closing in favor of #232.

@kevchoi
kev (kevchoi) deleted the kevchoi/20260928-autoevals-voice-scorers branch October 2, 2026 22:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant