Downloads · 30 days
18
7% of all-time downloads
MEscriva/gilbert-fr-source
gilbert-fr-source is a automatic speech recognition model from MEscriva. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as mit.
Gilbert-FR-Source is the foundational baseline model for the Gilbert research project, a comprehensive initiative focused on developing state-of-the-art automatic speech recognition (ASR) systems optimized for French…
Downloads · 30 days
18
7% of all-time downloads
All-time downloads
244
Public
Parameters
1.6B
13.6 GB on disk
Likes
2
Public
Click a slice to open those files.
.bin7.3 GB · 53%
From the Hugging Face model README
Gilbert-FR-Source is the foundational baseline model for the Gilbert research project, a comprehensive initiative focused on developing state-of-the-art automatic speech recognition (ASR) systems optimized for French language applications. This model serves as the frozen reference point for all subsequent research, fine-tuning, and development work within the Gilbert ecosystem.
Important Notice on Intellectual Property:
MEscriva/gilbert-fr-source) is distributed under the MIT License, allowing research and commercial use.The Gilbert project is a systematic research and development effort aimed at creating highly specialized ASR systems for:
This baseline model provides the controlled starting point for all experimental work, ensuring reproducibility and enabling fair comparison across different research directions.
This model is intended for:
While this baseline model can be used directly, production deployments should use specialized Gilbert models that are optimized for specific use cases and domains. Contact the Gilbert team for production-grade models.
The following WER (Word Error Rate) scores serve as baseline reference for future Gilbert model development:
| Dataset | WER | Notes |
|---|---|---|
| MLS (FR) | 3.98% | Multilingual LibriSpeech French |
| Common Voice FR (v13.0) | 7.28% | Diverse French speech |
| VoxPopuli (FR) | 8.91% | European Parliament speeches |
| Fleurs (FR) | 4.84% | FLORES evaluation |
| African Accented French | 4.20% | Regional accent evaluation |
Note: These results represent the upper bound before targeted fine-tuning. Future Gilbert variants will be evaluated against these baselines to measure improvement.
pip install transformers torch torchaudio librosa soundfile
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
import torch
model_id = "MEscriva/gilbert-fr-source"
device = "cuda" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.float16 if device == "cuda" else torch.float32
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForSpeechSeq2Seq.from_pretrained(
model_id,
torch_dtype=torch_dtype,
low_cpu_mem_usage=True
)
model.to(device)
# Process audio
audio_path = "your_audio.wav"
inputs = processor(audio_path, return_tensors="pt", sampling_rate=16000)
inputs = {k: v.to(device) for k, v in inputs.items()}
with torch.no_grad():
generated_ids = model.generate(
inputs["input_features"],
language="fr",
task="transcribe"
)
transcription = processor.batch_decode(
generated_ids,
skip_special_tokens=True
)[0]
import whisper
# Load the model
model = whisper.load_model("large-v3")
# Transcribe French audio
result = model.transcribe(
"audio.wav",
language="fr",
task="transcribe"
)
print(result["text"])
This model serves as:
This baseline model inherits known limitations from Whisper and the underlying training data:
Understanding and quantifying these limitations is a core objective of the Gilbert research roadmap.
The following specialized models will be developed as independent checkpoints from this baseline:
Gilbert-FR-Longform-v1
Gilbert-FR-Accents-v1
Gilbert-FR-Telephone-v1
Gilbert-Multilingual-v1
All future Gilbert models are the exclusive intellectual property of Lexia France and will include detailed evaluation reports adhering to research reproducibility standards.
This baseline model (MEscriva/gilbert-fr-source) is distributed under the MIT License, allowing:
See the LICENSE file for full terms.
Important: While this baseline model is available under MIT License:
For licensing inquiries regarding Gilbert project models, contact: mathis@lexiapro.fr
If you use this baseline model in your research, please cite:
@software{gilbert_fr_source_2024,
title={Gilbert-FR-Source: Research Baseline for French Automatic Speech Recognition},
author={MEscriva and Lexia France},
year={2024},
url={https://huggingface.co/MEscriva/gilbert-fr-source},
version={0.1},
note={Research baseline for the Gilbert project}
}
This baseline model is based on:
We acknowledge the contributions of the open-source community and the original Whisper research team.
For research collaboration, evaluation access, or technical inquiries:
© 2024 Lexia France. All rights reserved for Gilbert project derivatives.