Downloads · 30 days
8
15% of all-time downloads
OpenVoiceOS/ARK-ASR-0.6B-onnx
ARK-ASR-0.6B-onnx is a automatic speech recognition model from OpenVoiceOS. Use it when you need speech turned into text. It is set up for onnx-asr. The card lists the license as apache-2.0.
ONNX export of Audio8/ARK-ASR-0.6B for onnx-asr. All credit for the model goes to Audio8 (AutoArk AI). This repository only contains the converted graphs; the weights are the original ones.
Downloads · 30 days
8
15% of all-time downloads
All-time downloads
53
Public
Repo size
6.5 GB
Likes
0
Public
Click a slice to open those files.
.data5.2 GB · 80%
From the Hugging Face model README
ONNX export of Audio8/ARK-ASR-0.6B for onnx-asr. All credit for the model goes to Audio8 (AutoArk AI). This repository only contains the converted graphs; the weights are the original ones.
The model is a speech-LLM: a Whisper-large-v3-style audio encoder with rotary position embeddings, an MLP adapter that merges four encoder frames into one embedding, and a Qwen2 0.6B causal language model that writes the transcription.
pip install onnx-asr[cpu,hub]
import onnx_asr
model = onnx_asr.load_model("speech-llm", "OpenVoiceOS/ARK-ASR-0.6B-onnx")
print(model.recognize("audio.wav"))
# int8 weights
model = onnx_asr.load_model("speech-llm", "OpenVoiceOS/ARK-ASR-0.6B-onnx", quantization="int8")
| File | Contents |
|---|---|
encoder.onnx | audio encoder and MLP adapter, log-mel features in, LM embeddings out |
embed_tokens.onnx | token embedding table |
decoder.onnx | Qwen2 decoder with KV cache, logits out |
*.int8.onnx | dynamically quantized int8 weights |
config.json | model type, prompt token ids, suppressed token ids |
vocab.json | tokenizer vocabulary for detokenization |
| Graph | Inputs | Outputs |
|---|---|---|
encoder.onnx | input_features (1, 128, frames) | audio_embeds (1, frames/8, 896) |
embed_tokens.onnx | input_ids (1, S) | inputs_embeds (1, S, 896) |
decoder.onnx | inputs_embeds (1, S, 896), attn_bias (1, 1, S, P+S), position_ids (1, S), past_key_values.{0..23}.{key,value} (1, 2, P, 64) | logits (1, S, 163958), present.{0..23}.{key,value} (1, 2, P+S, 64) |
Four FLEURS clips (2 English, 2 Mandarin), greedy decoding, compared against the PyTorch model in float32:
Speed on a 12-core CPU under load: RTF 0.35 to 0.61 (fp32) and 0.09 to 0.29 (int8).
Apache 2.0, the same licence as the source model. The model was published by Audio8; see the source repository and the paper arXiv:2605.28139.