Downloads · 30 days
0
espnet/kosp2e-asr-ko
kosp2e-asr-ko is a automatic speech recognition model from espnet. Use it when you need speech turned into text. It is set up for espnet. The card lists the license as cc-by-4.0.
This is the ESPnet2 recipe for the KoSP2E (Korean Speech Perception and Production Experiment) dataset.
Downloads · 30 days
0
Access
Public
Updated Sep 19, 2026
Repo size
447 MB
Likes
0
Public
Click a slice to open those files.
.pth447 MB · 100%
From the Hugging Face model README
import librosa
from espnet2.bin.asr_inference import Speech2Text
speech2text = Speech2Text.from_pretrained(model_tag="espnet/kosp2e-asr-ko")
# librosa resamples and mixes to one channel, so any file works; 16000 is
# what nearly every espnet recogniser is trained on - check this model's
# config if its audio is not 16 kHz
speech, rate = librosa.load("audio.wav", sr=16000, mono=True)
text, *_ = speech2text(speech)[0]
print(text)
This is the ESPnet2 recipe for the KoSP2E (Korean Speech Perception and Production Experiment) dataset.
The KoSP2E dataset is a large-scale Korean speech corpus designed for speech perception and production experiments. This recipe provides a full ASR pipeline using ESPnet2 with both Transformer and Conformer architectures.
Environment
| dataset | Snt | Wrd | Corr | Sub | Del | Ins | Err | S.Err |
|---|---|---|---|---|---|---|---|---|
| test | 2320 | 22337 | 77.1 | 20.4 | 2.6 | 4.4 | 27.4 | 76.4 |
| dataset | Snt | Wrd | Corr | Sub | Del | Ins | Err | S.Err |
|---|---|---|---|---|---|---|---|---|
| test | 2320 | 84267 | 92.5 | 5.7 | 1.8 | 1.7 | 9.2 | 76.4 |
| dataset | Snt | Wrd | Corr | Sub | Del | Ins | Err | S.Err |
|---|---|---|---|---|---|---|---|---|
| test | 2320 | 65361 | 89.4 | 8.6 | 2.0 | 2.1 | 12.7 | 76.4 |