Downloads · 30 days
353
54% of all-time downloads
litert-community/whisper-medium
whisper-medium is a automatic speech recognition model from litert-community. Use it when you need speech turned into text. It is set up for litert. The card lists the license as mit.
Quantized TFLite (LiteRT) exports of openai/whisper-medium: 30 s fixed-window graphs split into encode and decode signatures, matching the graph interface of litert-community/whisper-tiny and litert-community/whisper-…
Downloads · 30 days
353
54% of all-time downloads
All-time downloads
653
Public
Repo size
1.5 GB
Likes
1
Public
Click a slice to open those files.
.tflite1.5 GB · 100%
From the Hugging Face model README
Quantized TFLite (LiteRT) exports of openai/whisper-medium: 30 s fixed-window graphs split into encode and decode signatures, matching the graph interface of litert-community/whisper-tiny and litert-community/whisper-base — drop-in for pipelines built against those graphs.
Exported by the LiteRT-LM-Unity project from the openai checkpoint (transformers TFWhisperForConditionalGeneration → two-signature 30 s graph), then post-training quantized. f32 source exports (~3 GB) are retained by the producing project and available on request.
| File | Size | Recipe |
|---|---|---|
whisper_medium_30s_i8.tflite | 794 MB | dynamic-range int8 weights (channelwise), fp32 activations |
whisper_medium_30s_i4.tflite | 634 MB | dynamic_wi4b64_afp32 mixed ("L1"): int4 blockwise-64 weights (fp16 scales) with the token-embedding/logits scopes kept at int8 channelwise |
tokenizer.json from openai/whisper-medium (differs from whisper-base's tokenizer).Quantizer: ai-edge-quantizer 0.8.0, post-training dynamic-range.
dynamic_wi8_afp32 scheme).dynamic_wi4b64_afp32 (int4 weights, blockwise-64, fp16 scales) with int8 overrides on the token-embedding/logits table. Pure full-scope int4 was tested and rejected (Korean transcription errors); an additional encoder-at-int8 variant produced identical transcripts and was discarded for size. int4 channelwise and blockwise-32 recipes are known-bad (quality collapse / immediate EOS) and were not used.9-clip Korean/English test set (Korean tactical-report sentence "2025년 3월 5일 전술평가 결과 보고", English equivalent, weather/status sentences, short Korean voice commands). Greedy decode, desktop CPU (XNNPACK, 8 threads). CER against punctuation-normalized references (space-removed); "exact" = normalized exact match.
| Variant | Exact /9 | CER ko | CER en | Avg RTF | Avg ms/step |
|---|---|---|---|---|---|
| i8 | 7 | 0.042 | 0.000 | 1.67 | 295 |
| i4 | 7 | 0.042 | 0.000 | 1.74 | 311 |
i4 reproduces i8's transcripts on every clip. The only content error in either tier is one short voice-command clip heard as 음향 증가 instead of 음량 증가 (CER 0.250 on that clip); all other misses vs exact are spacing-only or filename artifacts.
Validated with the LiteRT (ai-edge-litert) interpreter on Windows x86_64 CPU (XNNPACK) and within the LiteRT-LM-Unity v0.14.0 pipeline (LiteRT-LM v0.14.0, Windows x86_64 + Android arm64, Snapdragon 865-class device).
<|ko|> / <|en|>).LiteRT-LM-Unity (release v0.14.0-unity). Whisper weights are MIT (OpenAI); quantized derivatives inherit the base license.