Modernise local Whisper: transformers.js v4, WebGPU, windowed inference - #313
Merged
Merged
Conversation
Upgrade the whisper worker from @xenova/transformers 2.11.0 (WASM/CPU only) to @huggingface/transformers 4.2.0. Inference runs on WebGPU when available (fp32 encoder + q4 decoder, mirroring the official demos — fp16 decoders are numerically fragile and caused repetition-loop hallucinations) and falls back to WASM/q8 otherwise, including a rebuild-and-retry when a GPU passes init but fails at inference time. Long files are transcribed in 5-minute windows with 10s overlap, stitched at the inter-word gap nearest the overlap midpoint so seams land in silence. This bounds inference memory; together with transferring the decoded buffer to the worker (previously structured-cloned in full) and closing the decode AudioContext (previously leaked per run) it addresses the browser lock-up on large files. The model menu moves to the onnx-community *_timestamped exports (the v3+ replacement for the revision=output_attentions hack) with download sizes in the labels, adds Large v3 Turbo (q4f16) and drops distil-small (no timestamped export). The pipeline is cached per model and disposed on switch. Model-download and transcription progress are shown in the loader, and worker errors now render an error message instead of an eternal spinner. Null final-word timestamps no longer produce NaN data-d values. Closes #308
…rder
Some of the *_timestamped quantized exports predate the v4 ONNX runtime
and fail session creation ('Missing required scale ... MatMulNBits') —
base.en's q8 decoder among them. Try the preferred dtype first, then a
variant that loads everywhere (fp32, or q4 for turbo), per device.
Dispose the previous pipeline before loading the next model rather than
after, so two models are never resident at once, and stop chaining
.catch() onto dispose() — a non-promise return there poisoned every
device attempt and broke model switching entirely.
When the whisper decoder enters a repetition loop its word timestamps
collapse to a single value; within each run of same-start words keep
only the first occurrence of each word so hallucination loops can't
flood the transcript. Clamp data-d to >= 0 (zero-length words rendered
as data-d='-1') and surface the underlying error text on the error
screen.
With WebGPU, Base runs faster than Tiny did on the old CPU-only stack, and Tiny's failure mode on long or noisy audio is fabrication and repetition loops rather than graceful accuracy loss. Label Tiny as fastest/least accurate and steer long-form transcription to Base or larger in the help copy.
JSDoc @Version headers on the whisper worker and client module, app version comment and meta tag in index.html bumped to 0.6.6, and cache-busting query strings on the two changed whisper assets.
Transcribe-progress messages only arrive when a whole 5-minute window
completes, so between them the percentage sits still. A 1s ticker
appends elapsed time to the phase message ('Transcribing… 40% · 1m
23s') as a liveness signal, started on submission and stopped on
result or error.
transformers.js merges its internal 30s chunks by matching tokens across the seam; when the two decodes of the overlap disagree the merge can fail and a stretch of words is emitted twice (seen with the base model on a ~57s clip: 'Finally, export…' decoded once as 'audiovisual.' and again as 'audio visual media.'). The re-emit starts with word timestamps rewinding, which real speech never does — on a rewind of more than a second, truncate back to where the re-decode begins and let the later decode win.
Whisper marks word starts with a leading space on the chunk text; chunks without one are fragments of the previous word. The renderer appends a space to every chunk, so 'speech-to-text' came out as 'speech -to -text'. Merge fragments into their parent chunk, spanning its timing across the whole word.
Each span carried both whisper's leading space and the renderer's trailing space. Trimming also makes the capitalisation check examine the first letter rather than the space (which compared equal to its uppercase, so every word counted as capitalised and sentences incremented on any period), and guards against empty chunks.
This was referenced Jun 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #308
What this does
Replaces the local Whisper stack (
@xenova/transformers2.11.0, WASM/CPU only) with transformers.js 4.2.0, adding WebGPU inference, theonnx-community/*_timestampedmodel exports, model-download/transcription progress, windowed inference for long files, proper error handling, and a set of output-quality fixes found during testing.Worker (
js/whisper.worker.js, rewritten)*_timestampedquantized exports predate the v4 ONNX runtime and fail session creation (e.g. base.en's q8 decoder: "Missing required scale … MatMulNBits"); each device tries its preferred dtype then a variant that loads everywhere (fp32, or q4 for turbo).Output-quality fixes (worker)
All three exploit the same invariant — real speech only moves forward in time:
Client (
js/hyperaudio-lite-editor-whisper.js)AudioContextis closed after use (it leaked one per run) — together with windowing this addresses the browser lock-up on large files.NaN/negativedata-d.Model menu (
index.html)_timestampedexports with download sizes in the labels; added Large v3 Turbo (~560 MB, GPU recommended); dropped distil-small (no timestamped export).Versioning (0.6.6)
Per the project convention:
@version 0.6.6 — last changed in release 0.6.6JSDoc headers on both JS files, index.html version comment +<meta name="version">bumped, and?v=0.6.6cache-busting on the whisper script tag and worker path.Tested
--disable-features=WebGPU) and seam spot-checks at the 5-minute boundaries.Follow-ups filed
sanitise()crash surfaced by error screens🤖 Generated with Claude Code