Downloads · 30 days
46
19% of all-time downloads
BaseLayer/rus_uzb_eng_stt
rus_uzb_eng_stt is a automatic speech recognition model from BaseLayer. Use it when you need speech turned into text. The card lists the license as mit.
Downloads · 30 days
46
19% of all-time downloads
All-time downloads
237
Public
Parameters
764M
3.1 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors3.1 GB · 100%
From the Hugging Face model README
Ishlab chiqaruvchi: BaseLayer
Base model umumiy holatlarda yaxshi ishlaydi, lekin past ovozda (shivirlashga yaqin, xonadan uzoqdan yozilgan) nutqni tanishda xatoliklar ko'proq edi. Shu muammoni hal qilish uchun model qo'shimcha ma'lumot bilan qayta o'qitildi.
Whisper me'morchiligining tabiiy ko'p tillilik xususiyati tufayli, model faqat o'zbekcha emas — o'zbek, ingliz va rus tillarida, hattoki bir gap ichida tillar aralashib kelganda ham (code-switching, masalan "keyin meeting'da discuss qilamiz", "hisobotni project bo'yicha tayyorladim") nutqni ishonchli tarzda matnga aylantira oladi.
| Parametr | Qiymat |
|---|---|
| Learning rate | 1e-5 |
| Epochs | 5 (early stopping bilan 4-epochda to'xtatildi) |
| Batch size | 2 (gradient accumulation 8, effektiv batch 16) |
| Optimal checkpoint | Epoch 2 |
| Mixed precision | fp16 |
| Epoch | Training Loss | Validation Loss | WER |
|---|---|---|---|
| 1 | 0.1927 | 0.1058 | 11.18% |
| 2 (tanlangan) | 0.0465 | 0.1121 | 8.82% |
| 3 | 0.0287 | 0.1185 | 10.39% |
| 4 | 0.0178 | 0.1185 | 9.97% |
Eng yaxshi natija (WER 8.82%) 2-epochda qo'lga kiritildi va overfitting boshlanishidan oldin shu checkpoint yakuniy model sifatida tanlandi.
from transformers import WhisperProcessor, WhisperForConditionalGeneration
import librosa
processor = WhisperProcessor.from_pretrained("BaseLayer/uzbek_stt_model")
model = WhisperForConditionalGeneration.from_pretrained("BaseLayer/uzbek_stt_model")
audio, sr = librosa.load("audio.wav", sr=16000)
input_features = processor(audio, sampling_rate=16000, return_tensors="pt").input_features
predicted_ids = model.generate(input_features, language="uz", task="transcribe")
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
print(transcription)
Yoki pipeline orqali:
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="BaseLayer/sota_uzbek_stt_lowvolume",
chunk_length_s=30,
device="cuda"
)
result = pipe("audio.wav", generate_kwargs={"language": "uz", "task": "transcribe"})
print(result["text"])
MIT — bemalol foydalanish, o'zgartirish va tijoriy maqsadlarda qo'llash mumkin. Manba dataset (Beehzod/uzbek_speech_data) ham MIT litsenziyali.