Downloads · 30 days
4
17% of all-time downloads
OpenVoiceOS/paraformer-en-onnx
paraformer-en-onnx is a automatic speech recognition model from OpenVoiceOS. Use it when you need speech turned into text. It is set up for onnx-asr. The card lists the license as apache-2.0.
The English Paraformer for onnx-asr. The vocabulary is 4199 word pieces joined by the @@ subword marker. Output is lowercase and unpunctuated.
Downloads · 30 days
4
17% of all-time downloads
All-time downloads
24
Public
Repo size
1.1 GB
Likes
0
Public
Click a slice to open those files.
.onnx1.1 GB · 100%
From the Hugging Face model README
The English Paraformer for onnx-asr. The vocabulary is 4199 word pieces joined by the @@ subword marker. Output is lowercase and unpunctuated.
Paraformer is the Alibaba FunASR offline non-autoregressive recognizer. A SAN-M encoder reads the audio, a CIF predictor decides how many tokens the utterance has, and a single pass decoder emits all of them at once. There is no decoding loop, so one forward pass gives the transcript.
The ONNX graphs here are copied byte for byte from the
sherpa-onnx exports by
csukuangfj. Only the side files changed: tokens.txt became
vocab.txt, and config.json carries the FunASR frontend statistics from am.mvn.
The paraformer model type is on the feat/paraformer branch of the TigreGotico
onnx-asr fork.
pip install "onnx-asr[cpu,hub] @ git+https://github.com/TigreGotico/onnx-asr@feat/paraformer"
import onnx_asr
model = onnx_asr.load_model("OpenVoiceOS/paraformer-en-onnx")
print(model.recognize("audio.wav"))
| Item | Value |
|---|---|
| Input | speech, float32, [batch, num_frames, 560] |
| Input | speech_lengths, int32, [batch] |
| Output | logits, float32, [batch, num_tokens, 4200] |
| Output | token_num, int32, [batch], the CIF token count |
The 560 dim input is the FunASR frontend: an 80 dim kaldi fbank of a waveform scaled to
the int16 range, then a low frame rate stack of 7 frames with a hop of 6, then the
am.mvn mean variance statistics. onnx-asr computes the fbank with its wespeaker
preprocessor and applies the LFR stack and the CMVN in the runtime. Decoding is one
argmax per logits row, stopping at </s> and never reading past token_num.
A streaming Paraformer also exists upstream. It uses a different graph with encoder and decoder states and needs a streaming runtime, which onnx-asr does not have yet (upstream issue #21). Only the offline model is mirrored here.
iic/speech_paraformer_asr-en-16k-vocab4199-pytorch, Apache-2.0.model.onnx (856 MB) and model_int8.onnx (230 MB).
4 FLEURS clips, native FunASR on the same source checkpoint with dither = 0.
| Clip | fp32 | int8 |
|---|---|---|
| en_1 | identical | identical |
| en_2 | identical | identical |