Investigation and write-up done with Claude Code; the failures were reproduced and verified by me in real browsers.
Summary
The q8 (quantized) decoder of onnx-community/whisper-base.en_timestamped fails session creation under @huggingface/transformers 4.2.0's ONNX runtime:
Error: Can't create a session. ERROR_CODE: 1, ERROR_MESSAGE: qdq_actions.cc:137
TransposeDQWeightsForMatMulNBits Missing required scale:
model.decoder.embed_tokens.weight_merged_0_scale
for node: model.decoder.embed_tokens.weight_transposed_DequantizeLinear
Reproduced in both Chrome and Firefox, device: "wasm", dtype: "q8". The fp32 variant of the same model loads and runs fine, as does the WebGPU config ({encoder_model: "fp32", decoder_model_merged: "q4"}).
Repro
import { pipeline } from "@huggingface/transformers"; // 4.2.0
await pipeline("automatic-speech-recognition",
"onnx-community/whisper-base.en_timestamped",
{ device: "wasm", dtype: "q8" });
// -> Can't create a session ... Missing required scale ...
Likely cause
The *_timestamped exports predate the v4 runtime; its optimizer now rewrites the QDQ pattern into MatMulNBits (TransposeDQWeightsForMatMulNBits) and requires scale tensors the older q8 export doesn't carry.
Impact
q8 is the natural default for CPU/WASM users (and what the official demos use for the wasm path), so on machines without WebGPU this model is unusable at its intended dtype — apps either fail outright or must fall back to the ~292 MB fp32 download (what we do now in hyperaudio/hyperaudio-lite-editor#313).
Scope notes
- Verified broken:
whisper-base.en_timestamped q8 (decoder_model_merged_quantized.onnx).
- Not yet tested: the q8 variants of the other
*_timestamped exports (tiny/small/.en/multilingual, large-v3-turbo) — they're from the same export batch, so plausibly affected too.
Ask
Re-export the q8 (and possibly sibling) variants of the *_timestamped models against the current runtime — or, if re-export isn't planned, a note on the model cards steering wasm users to fp32 would save others the debugging session.
Summary
The
q8(quantized) decoder ofonnx-community/whisper-base.en_timestampedfails session creation under@huggingface/transformers4.2.0's ONNX runtime:Reproduced in both Chrome and Firefox,
device: "wasm",dtype: "q8". The fp32 variant of the same model loads and runs fine, as does the WebGPU config ({encoder_model: "fp32", decoder_model_merged: "q4"}).Repro
Likely cause
The
*_timestampedexports predate the v4 runtime; its optimizer now rewrites the QDQ pattern intoMatMulNBits(TransposeDQWeightsForMatMulNBits) and requires scale tensors the older q8 export doesn't carry.Impact
q8is the natural default for CPU/WASM users (and what the official demos use for the wasm path), so on machines without WebGPU this model is unusable at its intended dtype — apps either fail outright or must fall back to the ~292 MB fp32 download (what we do now in hyperaudio/hyperaudio-lite-editor#313).Scope notes
whisper-base.en_timestampedq8 (decoder_model_merged_quantized.onnx).*_timestampedexports (tiny/small/.en/multilingual, large-v3-turbo) — they're from the same export batch, so plausibly affected too.Ask
Re-export the q8 (and possibly sibling) variants of the
*_timestampedmodels against the current runtime — or, if re-export isn't planned, a note on the model cards steering wasm users to fp32 would save others the debugging session.