Downloads · 30 days
11
12% of all-time downloads
zeromodels/s2t-medium-librispeech-asr
s2t-medium-librispeech-asr is a automatic speech recognition model from zeromodels. Use it when you need speech turned into text. It is set up for zeromodels. The card lists the license as mit.
[](https://github.com/IMvision12/ZeroModels) [](https://imvision12.github.io/ZeroModels/speech2text/) [](https://huggingface.co/collections/zeromodels/speech2text-6a8eaf32c86fd39eeacbb3d0)
Downloads · 30 days
11
12% of all-time downloads
All-time downloads
90
Public
Repo size
321 MB
Likes
1
Public
Click a slice to open those files.
.h5320 MB · 100%
From the Hugging Face model README
Paper: fairseq S2T: Fast Speech-to-Text Modeling with fairseq (arXiv:2010.05171) · HF Papers
Speech2Text (fairseq S2T) is a classic encoder-decoder ASR model trained on LibriSpeech. Transcripts are lowercase and unpunctuated, matching the training label style (unlike Whisper / Moonshine casing).
For more details on the model, please go to the upstream model card.
Pure-Keras 3 conversion of facebook/s2t-medium-librispeech-asr for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.
This is an ASR checkpoint (Speech2TextConditionalGenerate, medium).
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
import soundfile as sf
from zeromodels.models.speech2text import (
Speech2TextProcessor,
Speech2TextConditionalGenerate,
)
model = Speech2TextConditionalGenerate.from_weights("zeromodels/s2t-medium-librispeech-asr")
processor = Speech2TextProcessor.from_weights("zeromodels/s2t-medium-librispeech-asr")
audio, sr = sf.read("your_audio.wav", dtype="float32") # 16 kHz mono
text = model.generate(audio, processor)
print(repr(text[0])) # lowercase, unpunctuated LibriSpeech style
Load any Speech2Text variant the same way with from_weights("zeromodels/<variant>"):
| Variant | Hub |
|---|---|
s2t-small-librispeech-asr | zeromodels/s2t-small-librispeech-asr |
s2t-medium-librispeech-asr | zeromodels/s2t-medium-librispeech-asr |
s2t-large-librispeech-asr | zeromodels/s2t-large-librispeech-asr |
KERAS_BACKEND before importing Keras / zeromodels.Speech2TextProcessor.from_weights(...) so fbank settings match.hf: prefix, e.g. Speech2TextConditionalGenerate.from_weights("hf:facebook/s2t-medium-librispeech-asr").A huge thank you to the Facebook fairseq S2T authors for creating and releasing these models.
License: MIT.