Downloads · 30 days
0
espnet/owls_1b_45k
owls_1b_45k is a automatic speech recognition model from espnet. Use it when you need speech turned into text.
Downloads · 30 days
0
Access
Public
Updated Sep 19, 2026
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.pth2.2 GB · 100%
From the Hugging Face model README
import librosa
from espnet2.bin.s2t_inference import Speech2Text
s2t = Speech2Text.from_pretrained(
model_tag="espnet/owls_1b_45k", lang_sym="<eng>", task_sym="<asr>", beam_size=5
)
# OWSM is trained on 16 kHz; each call decodes 30 s, padded or trimmed
speech, rate = librosa.load("audio.wav", sr=16000, mono=True)
text, token, token_int, text_nospecial, hyp = s2t(speech)[0]
print(text_nospecial) # `text` keeps OWSM's own <eng><asr> markers
# for a recording longer than 30 s: s2t.decode_long(speech) -> (start, end, text)