Downloads · 30 days
88
6% of all-time downloads
aihpi/FrWhisper
FrWhisper is a automatic speech recognition model from aihpi. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
FrWhisper is a fine-tune of openai/whisper-large-v3 for spoken French, optimised for conversational speech with disfluencies (hesitations such as "euh", repetitions, interjections, spoken numbers). It is trained on ma…
Downloads · 30 days
88
6% of all-time downloads
All-time downloads
1.5K
Public
Parameters
1.5B
12.3 GB on disk
Likes
11
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
FrWhisper is a fine-tune of openai/whisper-large-v3 for spoken French, optimised for conversational speech with disfluencies (hesitations such as "euh", repetitions, interjections, spoken numbers). It is trained on material from the ESLO and LangAge corpora.
This repository keeps both releases as git tags:
v2 (this version, main): trained on an improved subset of ESLO/LangAge data.v1: the earlier model, trained on a less optimal subset of ESLO/LAngAge data. Load it with revision="v1".from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="aihpi/FrWhisper") # v2 (latest)
asr_v1 = pipeline("automatic-speech-recognition", model="aihpi/FrWhisper", revision="v1")
WER / CER on a fixed 1000-sample held-out subset of the corrected test split (greedy decoding, forced French, lowercase + punctuation-stripped normalisation):
| Model | WER (all) | WER ESLO | WER LangAge | CER (all) |
|---|---|---|---|---|
| Whisper large-v3 (base) | 50.81 | 48.11 | 54.64 | 35.44 |
| FrWhisper v1 (pre-improvement) | 34.13 | 27.74 | 43.19 | 23.68 |
| FrWhisper v2 (this model) | 29.87 | 22.41 | 40.46 | 20.95 |
v2 improves over the base model by ~21 WER points and over v1 by ~4 points overall, with the largest gain on ESLO (the corpus most affected by the improved subset).
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="aihpi/FrWhisper")
result = asr("audio.wav") # 16 kHz mono; French is forced by the generation config
print(result["text"])
For long files, enable chunking: pipeline(..., chunk_length_s=30).
Released under CC BY-NC-SA 4.0, subject to the terms of the underlying ESLO and LangAge corpora. ESLO material is Copyright (c) 2012 Université d'Orléans / LLL, freely available for non-commercial use under a Creative Commons licence.
@misc{frwhisper2025,
title={FrWhisper: Whisper Large-v3 fine-tuned for French conversational speech},
author={Hanno Müller, Annette Gerstenberg},
year={2025},
note={Fine-tuned on the LangAge and ESLO corpora}
}
The AI Service Centre Berlin Brandenburg is funded by the Federal Ministry of Research, Technology and Space under the funding code 01IS22092.
For questions about this model, please open an issue in the repository or contact [email protected].