Sitelet https://github.com/hyperaudio/hyperaudio-lite-editor/pull/313
Skip to content

Modernise local Whisper: transformers.js v4, WebGPU, windowed inference - #313

Merged
maboa merged 9 commits into
mainfrom
308-modernise-local-whisper
Jun 11, 2026
Merged

maboa merged 9 commits into
mainfrom
308-modernise-local-whisper

Conversation

@maboa

@maboa maboa commented Jun 10, 2026 •

Copy link
Copy Markdown
Member

Closes #308

What this does

Replaces the local Whisper stack (@xenova/transformers 2.11.0, WASM/CPU only) with transformers.js 4.2.0, adding WebGPU inference, the onnx-community/*_timestamped model exports, model-download/transcription progress, windowed inference for long files, proper error handling, and a set of output-quality fixes found during testing.

Worker (js/whisper.worker.js, rewritten)

  • WebGPU first, WASM fallback — including a rebuild-and-retry on WASM if a GPU passes init but fails during inference. The chosen device/dtype is logged to the console.
  • Per-device dtypes matching the official demos: fp32 encoder + q4 decoder on WebGPU (fp16 decoders are numerically fragile — caused repetition-loop hallucinations in testing), q8 on WASM, q4f16 for Large v3 Turbo.
  • Dtype fallback: some *_timestamped quantized exports predate the v4 ONNX runtime and fail session creation (e.g. base.en's q8 decoder: "Missing required scale … MatMulNBits"); each device tries its preferred dtype then a variant that loads everywhere (fp32, or q4 for turbo).
  • Windowed inference: 5-minute windows with 10 s overlap, stitched at the inter-word gap nearest the overlap midpoint so seams land in silence, never mid-word. Bounds memory on hour-plus files.
  • Pipeline cached per model and disposed before the next model loads (switching previously left both resident).

Output-quality fixes (worker)

All three exploit the same invariant — real speech only moves forward in time:

  • Repetition-loop containment: when Whisper's decoder loops, its word timestamps collapse to a single value; within each run of same-start words only the first occurrence of each word is kept, so hallucination loops can't flood the transcript.
  • Rewind dedupe: transformers.js merges its internal 30 s chunks by token matching; when the two decodes of an overlap disagree (e.g. "audiovisual." vs "audio visual media.") the merge fails and re-emits a stretch of words. The re-emit starts with timestamps rewinding — truncate back to the rewind point and let the later decode (the one not truncated by a chunk edge) win.
  • Word-fragment merge: whisper marks word starts with a leading space; fragments without one are parts of the previous word, so "speech-to-text" now renders as one timed span instead of "speech -to -text".

Client (js/hyperaudio-lite-editor-whisper.js)

  • Loader live-updates with phase, percentage and elapsed time: "Preparing model… (4s)" → "Downloading model… 63% (12s)" → "Transcribing… 40% (1m 23s)". Progress only arrives per completed window, so the ticking clock is the liveness signal in between.
  • The decoded audio buffer is transferred to the worker instead of structured-cloned, and the decode AudioContext is closed after use (it leaked one per run) — together with windowing this addresses the browser lock-up on large files.
  • Worker errors and crashes render an error screen with the underlying message instead of an eternal spinner.
  • Whisper's leading word-boundary space is trimmed when rendering, so spans carry a single trailing space — this also makes the sentence/paragraph capitalisation check examine the first letter rather than the space (which always passed), giving better-placed paragraph breaks.
  • Null final-word timestamps no longer produce NaN/negative data-d.

Model menu (index.html)

  • Switched to the _timestamped exports with download sizes in the labels; added Large v3 Turbo (~560 MB, GPU recommended); dropped distil-small (no timestamped export).
  • Default is now Base English — with WebGPU it's faster than Tiny was on CPU, and Tiny's failure mode on long/noisy audio is fabrication rather than graceful degradation. Tiny remains, labelled "fastest, least accurate", with help copy steering long-form audio to Base+.

Versioning (0.6.6)

Per the project convention: @version 0.6.6 — last changed in release 0.6.6 JSDoc headers on both JS files, index.html version comment + <meta name="version"> bumped, and ?v=0.6.6 cache-busting on the whisper script tag and worker path.

Tested

  • WebGPU path verified on Tiny and Base, including the dtype fallback (exercised by base.en's broken q8 export) and mid-session model switching.
  • Hour-long file transcribed; a Tiny repetition loop was contained by the collapse logic and resolved by Base.
  • Rewind dedupe and fragment merge verified against the real-world failures that motivated them (~57 s clip, base model, wav and mp3).
  • Not formally exercised: a forced-WASM run (--disable-features=WebGPU) and seam spot-checks at the 5-minute boundaries.

Follow-ups filed

🤖 Generated with Claude Code

maboa added 9 commits June 10, 2026 20:58
Upgrade the whisper worker from @xenova/transformers 2.11.0 (WASM/CPU
only) to @huggingface/transformers 4.2.0. Inference runs on WebGPU when
available (fp32 encoder + q4 decoder, mirroring the official demos —
fp16 decoders are numerically fragile and caused repetition-loop
hallucinations) and falls back to WASM/q8 otherwise, including a
rebuild-and-retry when a GPU passes init but fails at inference time.

Long files are transcribed in 5-minute windows with 10s overlap,
stitched at the inter-word gap nearest the overlap midpoint so seams
land in silence. This bounds inference memory; together with
transferring the decoded buffer to the worker (previously
structured-cloned in full) and closing the decode AudioContext
(previously leaked per run) it addresses the browser lock-up on large
files.

The model menu moves to the onnx-community *_timestamped exports (the
v3+ replacement for the revision=output_attentions hack) with download
sizes in the labels, adds Large v3 Turbo (q4f16) and drops distil-small
(no timestamped export). The pipeline is cached per model and disposed
on switch. Model-download and transcription progress are shown in the
loader, and worker errors now render an error message instead of an
eternal spinner. Null final-word timestamps no longer produce NaN
data-d values.

Closes #308
…rder

Some of the *_timestamped quantized exports predate the v4 ONNX runtime
and fail session creation ('Missing required scale ... MatMulNBits') —
base.en's q8 decoder among them. Try the preferred dtype first, then a
variant that loads everywhere (fp32, or q4 for turbo), per device.

Dispose the previous pipeline before loading the next model rather than
after, so two models are never resident at once, and stop chaining
.catch() onto dispose() — a non-promise return there poisoned every
device attempt and broke model switching entirely.

When the whisper decoder enters a repetition loop its word timestamps
collapse to a single value; within each run of same-start words keep
only the first occurrence of each word so hallucination loops can't
flood the transcript. Clamp data-d to >= 0 (zero-length words rendered
as data-d='-1') and surface the underlying error text on the error
screen.
With WebGPU, Base runs faster than Tiny did on the old CPU-only stack,
and Tiny's failure mode on long or noisy audio is fabrication and
repetition loops rather than graceful accuracy loss. Label Tiny as
fastest/least accurate and steer long-form transcription to Base or
larger in the help copy.
JSDoc @Version headers on the whisper worker and client module, app
version comment and meta tag in index.html bumped to 0.6.6, and
cache-busting query strings on the two changed whisper assets.
Transcribe-progress messages only arrive when a whole 5-minute window
completes, so between them the percentage sits still. A 1s ticker
appends elapsed time to the phase message ('Transcribing… 40% · 1m
23s') as a liveness signal, started on submission and stopped on
result or error.
transformers.js merges its internal 30s chunks by matching tokens
across the seam; when the two decodes of the overlap disagree the merge
can fail and a stretch of words is emitted twice (seen with the base
model on a ~57s clip: 'Finally, export…' decoded once as 'audiovisual.'
and again as 'audio visual media.'). The re-emit starts with word
timestamps rewinding, which real speech never does — on a rewind of
more than a second, truncate back to where the re-decode begins and let
the later decode win.
Whisper marks word starts with a leading space on the chunk text;
chunks without one are fragments of the previous word. The renderer
appends a space to every chunk, so 'speech-to-text' came out as
'speech -to -text'. Merge fragments into their parent chunk, spanning
its timing across the whole word.
Each span carried both whisper's leading space and the renderer's
trailing space. Trimming also makes the capitalisation check examine
the first letter rather than the space (which compared equal to its
uppercase, so every word counted as capitalised and sentences
incremented on any period), and guards against empty chunks.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Modernise Whisper (Local): transformers.js v3 + WebGPU, newer models, chunked decoding to lift the file-size cap

1 participant