Downloads · 30 days
0
femustafa/voicedictation-models
voicedictation-models is a automatic speech recognition model from femustafa. Use it when you need speech turned into text. The card lists the license as mit.
Model artifacts for a privacy-focused, on-device dictation Android app. This repo hosts three model families:
Downloads · 30 days
0
Access
Public
Updated Sep 12, 2026
Repo size
3.9 GB
Likes
0
Public
Click a slice to open those files.
.onnx3.3 GB · 84%
From the Hugging Face model README
Model artifacts for a privacy-focused, on-device dictation Android app. This repo hosts three model families:
.bin),.onnx), and.onnx), built for pinned-language
(ur/PK) decoding on device.whisper.cpp (GGML) conversions of the
cheetos18/whisper-small-roman-urdu
fine-tune, consumed by the app's whisper.cpp backend.
The upstream repo hosts the Transformers/safetensors checkpoint; this repo hosts
the same model converted to the GGML .bin format that whisper.cpp consumes, so
the app can download it directly at runtime.
| File | Format | Size | Notes |
|---|---|---|---|
ggml-model-q4_0.bin | q4_0 quantized | ~139 MB | Default: fastest on device |
ggml-model-f16.bin | f16 | ~466 MB | Unquantized: higher quality |
Model is small (n_vocab 51865), multilingual.
whisper-cli -m ggml-model-q4_0.bin -f audio.wav -l auto
Important: always run with language auto-detect (-l auto), NEVER force
-l ur. Forcing the native language token leaks native-script (Urdu-script)
output; -l auto keeps the transcript in clean Roman/Latin script.
cheetos18/whisper-small-roman-urdu (model.safetensors, verified)whisper-quantize ggml-model.bin ggml-model-q4_0.bin q4_0aap ka kya haal hai?
under -l auto.Sherpa-onnx–format ONNX exports of the same
cheetos18/whisper-small-roman-urdu
fine-tune, consumed by the app's sherpa-onnx backend
(OfflineRecognizer.from_whisper). Two precision tiers share one token file.
| File | Tier | Size | Notes |
|---|---|---|---|
roman-urdu-encoder.onnx | fp32 | ~409 MB | whisper audio encoder |
roman-urdu-decoder.onnx | fp32 | ~559 MB | whisper text decoder (cross-attention KV caches) |
roman-urdu-encoder.int8.onnx | int8 | ~112 MB | MatMul weights int8, activations fp32 |
roman-urdu-decoder.int8.onnx | int8 | ~262 MB | MatMul weights int8, activations fp32 |
roman-urdu-tokens.txt | — | ~0.8 MB | token ID → token list (shared) |
Tier totals: fp32 ≈ 969 MB, int8 ≈ 375 MB.
Load with sherpa-onnx using the ONNX (whisper) recognizer and the shared token file, e.g.:
sherpa-onnx-offline \
--whisper-encoder=roman-urdu-encoder.onnx \
--whisper-decoder=roman-urdu-decoder.onnx \
--tokens=roman-urdu-tokens.txt \
audio.wav
Important: always decode with fixed language en (never ur). Forcing the
native-language token leaks Urdu-script output; English keeps the transcript in clean
Roman/Latin script.
cheetos18/whisper-small-roman-urdu (model.safetensors)quantize_dynamic(..., op_types_to_quantize=["MatMul"], weight_type=QInt8)).aap ka kya haal hai?; int8 as
aap ka kya haal hai.On-device Urdu recognition that honors an explicit ur/PK language pin —
unlike Dolphin's CTC model, which auto-detects and can leak into the wrong
script (e.g. Devanagari for Urdu speech). These are the ONNX encoder+decoder
pairs used by the app's onnxruntime-based Dolphin attention engine.
dolphin-attn/ (fp16/arm tier)
units.txt token ID → subword list (shared)
bpe.model sentencepiece BPE model (shared; DecodePieces)
base/
encoder.onnx fp16 arm, fused STFT+mel+CMVN in-graph, raw int16 in
decoder.onnx fp16 arm, attention decoder with FULL-logits output
small/
encoder.onnx
decoder.onnx
dolphin-attn-int8/ (int8 tier — see "int8 tier" below)
base/
encoder.onnx
decoder.onnx
units.txt
small/
encoder.onnx
decoder.onnx
units.txt
(The standalone units.txt at repo root is the same Dolphin subword vocab.)
| Tier | Model | encoder | decoder | total | notes |
|---|---|---|---|---|---|
| fp16 | base | ~124 MB | ~126 MB | ~250 MB | 6 layers, 8 heads; lower latency |
| fp16 | small | ~378 MB | ~322 MB | ~699 MB | 12 layers, 12 heads; higher quality |
| int8 | base | ~110 MB | ~153 MB | ~263 MB | int8 MatMul + output layer; see below |
| int8 | small | ~367 MB | ~381 MB | ~748 MB | int8 MatMul + output layer; see below |
decoder.onnx is graph-surgeried: it exposes /output_layer/Gemm_output_0
(the full logits over the whole 40,002-token vocabulary) in addition to max_logit_id,
so attention beam search can read the real distribution. Stock DakeQQ decoders
only output max_logit_id (argmax), which is insufficient for beam search.language_start/language_end-sliced logits output, but that is only for LID
(auto-detection); beam-searching it yields language-token garbage.attention (beam search), length-normalized, never
attention_rescoring or greedy — the rescoring/greedy paths override the ur/PK
pin and flip pinned Urdu to Devanagari.<sos> <ur> <PK> <asr> <notimestamp> and pass
language_start=136 / language_end=269.Where dolphin-attn/ is the upstream fp16/arm conversion (+ graph surgery),
dolphin-attn-int8/ was produced locally in this project because the upstream
stock int8 tier cannot satisfy the beam-search requirement:
.onnx), fp16/arm (.onnx),
and int8 (.ort). The .ort int8 files are onnxruntime flatbuffers; the
full-logits graph surgery operates on the ONNX proto and cannot modify them — so the
stock int8 tier can never expose /output_layer/Gemm_output_0 for beam search..onnx files are therefore the fp32 tier run through dynamic int8
quantization locally, followed by the same full-logits surgery as the fp16 tier:
quantize_dynamic(..., op_types_to_quantize=["MatMul"], weight_type=QInt8).Gather,
which dynamic quantization does not touch) stays fp32 — that one ~123 MB fp32 table
is why int8-small (~748 MB) is larger than fp16-small (~699 MB).onnx-community/dataocean-dolphin-asr
(fp16/arm tier)..ort and not graph-surgeriable.