Downloads · 30 days
139
13% of all-time downloads
Workmind/moonshine-streaming-small-ONNX
moonshine-streaming-small-ONNX is a automatic speech recognition model from Workmind. Use it when you need speech turned into text. It is set up for transformers.js. The card lists the license as mit.
ONNX export of UsefulSensors/moonshine-streaming-small (Moonshine v2, the moonshinestreaming architecture) for in-browser speech-to-text with transformers.js. No official ONNX export of this architecture exists upstre…
Downloads · 30 days
139
13% of all-time downloads
All-time downloads
1.1K
Public
Repo size
909 MB
Likes
0
Public
Click a slice to open those files.
.onnx909 MB · 100%
From the Hugging Face model README
ONNX export of UsefulSensors/moonshine-streaming-small (Moonshine v2, the moonshine_streaming architecture) for in-browser speech-to-text with transformers.js. No official ONNX export of this architecture exists upstream; this one was built and validated by Workmind. The export is batch/full-utterance (encoder + merged decoder with past-KV), not chunked-streaming.
import { pipeline } from '@huggingface/transformers';
const transcriber = await pipeline(
'automatic-speech-recognition',
'Workmind/moonshine-streaming-small-ONNX',
{ dtype: { encoder_model: 'q8', decoder_model_merged: 'q8' } },
);
const { text } = await transcriber(audio); // Float32Array, 16 kHz mono
Input length caveat: the v2 audio frontend reshapes input into 5 ms frames (80 samples @ 16 kHz) and fails on partial frames. transformers.js's generic feature extractor does not pad for you, so zero-pad each utterance to a multiple of 80 samples before calling the pipeline.
config.json says moonshine)transformers.js has no moonshine_streaming registration yet, so the published config.json intentionally declares model_type: "moonshine" with v1-shaped normalized-config keys. The stock transformers.js bundle then loads this model through its registered Moonshine v1 path; all v2 architectural differences (including the sliding-window encoder attention mask, synthesized all-ones inside the graph) are baked into the ONNX graphs, which are v1-compatible in I/O. Once native moonshine_streaming support ships upstream, a new revision of this repo will carry the real config.
| Variant | Encoder | Decoder (merged) | Notes |
|---|---|---|---|
| fp32 | encoder_model.onnx (196 MB) | decoder_model_merged.onnx (340 MB) | reference |
| q8 | encoder_model_quantized.onnx (71 MB) | decoder_model_merged_quantized.onnx (87 MB) | recommended; transcript-lossless vs fp32 |
| q4 | encoder_model_q4.onnx (56 MB) | decoder_model_merged_q4.onnx (116 MB) | transcript-lossless vs fp32 |
The q8 quantization keeps the causal-Conv audio frontend weights in fp32 (int8 Conv quantization silently produces empty transcripts).
fp32 greedy decode via onnxruntime matches the PyTorch reference token-for-token at the pinned revisions. In-browser (transformers.js 4.2.0 / onnxruntime-web WASM, M-series MacBook, 11 s clip): q8/q8 loads in ~2.3 s and transcribes in ~0.9 s with an exact transcript match against the Python reference.
Built from UsefulSensors/moonshine-streaming-small at revision 2c036506f23a09c18df5a50057599ba6d9280999 with transformers 5.14.1 and a patched optimum-onnx export config. The full reproducible pipeline (export, quantization, parity checks, config rewrite) lives at WorkmindAI/stt-models.
MIT, same as the upstream Moonshine weights by UsefulSensors / Moonshine AI. This repo is a converted/quantized redistribution of those weights with a compatibility-rewritten config; all credit for the model itself goes to the original authors.