Downloads · 30 days
36
100% of all-time downloads
wrice/whisper-tiny-grpo
whisper-tiny-grpo is a automatic speech recognition model from wrice. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as mit.
openai/whisper-tiny finetuned with GRPO (Group Relative Policy Optimization) using word error rate as the reward, across 81 Common Voice locales (79 distinct Whisper languages).
Downloads · 30 days
36
100% of all-time downloads
All-time downloads
36
Public
Parameters
37.8M
151 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors151 MB · 97%
From the Hugging Face model README
openai/whisper-tiny finetuned with GRPO (Group Relative Policy Optimization)
using word error rate as the reward, across 81 Common Voice locales (79 distinct Whisper languages).
Instead of cross-entropy against a single reference, the model samples a group of candidate transcriptions per clip, scores each by its WER against the ground truth, and is nudged toward the lower-error candidates with a policy-gradient objective regularized by a KL penalty to the original model. Training code: will-rice/whisper-rl.
Evaluated on the Common Voice 26.0 test split — 4,800 clips round-robin across all 81 locales in the index, greedy decoding, corpus-level rates.
| WER | CER | |
|---|---|---|
openai/whisper-tiny | 0.9718 | 0.5836 |
| this model | 0.6226 | 0.2789 |
| change | -0.3492 | -0.3047 |
Improves 69 of 81 locales.
Rates are bucketed by Common Voice locale, not by Whisper language. The two differ: zh-CN, zh-HK and zh-TW are separate locales that all prompt with the same <|zh|> token, and regional codes like pa-IN, ne-NP and sv-SE map to <|pa|>, <|ne|> and <|sv|>. 81 locales span 79 Whisper languages.
| locale | baseline WER | this model | change |
|---|---|---|---|
| kk | 2.164 | 0.686 | -1.479 |
| tk | 1.962 | 0.845 | -1.117 |
| pa-IN | 1.450 | 0.420 | -1.029 |
| fa | 1.612 | 0.791 | -0.821 |
| tt | 1.373 | 0.600 | -0.773 |
| hi | 0.996 | 0.244 | -0.751 |
| ka | 1.538 | 0.798 | -0.740 |
| sd | 1.794 | 1.055 | -0.738 |
| uz | 1.323 | 0.655 | -0.667 |
| ne-NP | 1.099 | 0.452 | -0.646 |
Baseline WER above 1.0 means the model emitted more tokens than the reference contains — runaway insertion, not merely inaccurate transcription. The largest gains are where that behaviour was worst.
| locale | baseline WER | this model | change |
|---|---|---|---|
| ko | 0.723 | 1.166 | +0.442 |
| ms | 0.665 | 0.981 | +0.316 |
| ar | 0.741 | 0.949 | +0.209 |
| en | 0.308 | 0.412 | +0.105 |
| vi | 0.614 | 0.675 | +0.061 |
| ru | 0.389 | 0.439 | +0.050 |
| he | 0.689 | 0.710 | +0.021 |
| sv-SE | 0.618 | 0.636 | +0.018 |
The regressions cluster in higher-resource languages where whisper-tiny was
already reasonable. The method trades some high-resource accuracy for large
low-resource gains.
Held-out in both senses that matter:
train with
CV26 test over six locales found 1 shared speaker in 4,900.test, which nothing in training touched.Baseline and finetuned models score the same clips in the same order — the
evaluation slice is materialized with take() and never shuffled — so a
difference can only come from the model.
openai/whisper-tiny, not against supervised finetuning
on the same data. This shows GRPO beats the base model, not that it beats SFT.ja, th, zh-*).from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("wrice/whisper-tiny-grpo")
processor = WhisperProcessor.from_pretrained("wrice/whisper-tiny-grpo")
The language token must be pinned at generation time; Whisper's own detection mislabels lower-resource clips, which is what the training setup assumes.