Downloads · 30 days
31
15% of all-time downloads
orgilj/moonshine-mn
moonshine-mn is a automatic speech recognition model from orgilj. Use it when you need speech turned into text. The card lists the license as apache-2.0.
Fine-tuned UsefulSensors/moonshine-base on Mongolian (Cyrillic) speech from Mozilla Common Voice.
Downloads · 30 days
31
15% of all-time downloads
All-time downloads
209
Public
Parameters
48.7M
195 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors195 MB · 100%
From the Hugging Face model README
Fine-tuned UsefulSensors/moonshine-base on Mongolian (Cyrillic) speech from Mozilla Common Voice.
| Checkpoint | WER |
|---|---|
| final (step 15000) | 11.88% |
import torch, librosa
from transformers import MoonshineForConditionalGeneration, AutoFeatureExtractor
from mn_tokenizer import MnBPETokenizer
from huggingface_hub import hf_hub_download
model = MoonshineForConditionalGeneration.from_pretrained("orgilj/moonshine-mn").eval()
fe = AutoFeatureExtractor.from_pretrained("orgilj/moonshine-mn")
tok = MnBPETokenizer(vocab_file=hf_hub_download("orgilj/moonshine-mn", "mn_bpe.model"))
def transcribe(path, num_beams=5):
audio, _ = librosa.load(path, sr=16000)
inp = fe(audio, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
ids = model.generate(
inp.input_values,
num_beams=num_beams,
max_new_tokens=180, # under the model's max_length=194
)
return tok.decode_ids(ids[0].tolist())
if __name__ == "__main__":
print(transcribe("/workspace/data/cv-corpus-24.0-2025-12-05/mn/clips/common_voice_mn_44590402.mp3"))
# From finetune-moonshine-asr repo:
python scripts/stream_mn.py --model orgilj/moonshine-mn --live