Downloads · 30 days
709
29% of all-time downloads
KrorngAI/TrorYongASR-small
TrorYongASR-small is a automatic speech recognition model from KrorngAI. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as other.
<div align="center" <picture <img src="figures/krorngai.png" width="30%" alt="KrorngAI" </picture </div <hr <div align="center" style="line-height:1" <a href="https://www.kimi.com" target="blank"<img alt="Chat" src="h…
Downloads · 30 days
709
29% of all-time downloads
All-time downloads
2.5K
Public
Parameters
134M
1.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors536 MB · 99%
From the Hugging Face model README
[!Note] This repository contains model weights and configuration files for the pre-trained model.
TrorYongASR is an Encoder-Decoder model for Automatic Speech Recognition (ASR) task. It is inspired by PARSeq and Whisper: the auditory-lingual decoder has only one transformer block.
<div align="center"> <picture> <img src="figures/architecture.png" width="100%" alt="TrorYongASR"> </picture> </div>TrorYongASR has 2 configurations:
<div align="center">| Model Size | Tiny | Small |
|---|---|---|
| Parameters | 29M | 135M |
| Audio Encoder | 4 layers, 6 heads | 12 layers, 12 heads |
| Text Decoder | 1 layer, 12 heads | 1 layer, 24 heads |
| Embedding Dim | 384 | 768 |
| Audio Context | 1500 | 1500 |
| Text Context | 1024 | 1024 |
Note: The audio array are processed to log-mel spectrogram with 80 mels (the same as Whisper models of the same size)
The evaluation assesses two capabilities — language detection and transcription — on two datasets (google/fleurs for Khmer and openslr/librispeech_asr for English). All results are from the test split of each dataset, representing the model's generalization ability to unseen data.
| Dataset | Language | Testing examples | Description |
|---|---|---|---|
| google/fleurs | Khmer | 765 | Multi-lingual dataset with Khmer language samples |
| librispeech.clean | English | 2620 | Clean speech dataset for English transcription |
Note: Audios longer than 30 seconds are excluded from the evaluation (that is why google/fleurs has 765 examples instead of 771).
Language detection measures model’s capability to recognize the spoken language from audio input. Since TrorYongASR currently supports 2 languages, this task becomes binary classification task. Classic metrics are used:
Results:
<div align="center">| Model | Metrics | Khmer (fleurs) | English (librispeech.clean) |
|---|---|---|---|
| Tiny | Precision | 100% | 100% |
| Recall | 100% | 100% | |
| F1-score | 100% | 100% | |
| Small | Precision | 100% | 99% |
| Recall | 96% | 100% | |
| F1-score | 98% | 99% |
Tiny size achieved perfect language detection performance on both datasets, indicating excellent binary classification capability for distinguishing between Khmer and English audio. Small size performs slightly worst by tending to predict English language.
The 100% language detection scores may appear unusually high. This is expected because during pre-training, the model performs permutations on word tokens starting from position 3, while the first three positions (start token, language token, and task token) remain fixed. Since language detection relies on the language token at position 1, and this token is never permuted during pre-training, the model can achieve perfect accuracy on language detection tasks.
For transcription task, 3 metrics below are used
Token Error Rate (TER) measures model's capability in predicting the next token given the audio input and the current sequence of tokens. This metric is weaker than Word Error Rate (WER) and Character Error Rate (CER) because it doesn't account for insertions, deletions, substitutions, and autoregression as comprehensively. Token Error Rate is used here because Khmer text lacks word boundaries, making WER and CER calculations challenging without additional preprocessing.
Transcription Results:
<div align="center">| Model | Metric | Khmer (fleurs) | English (librispeech.clean) | Mixed (Khmer + English) |
|---|---|---|---|---|
| Tiny | WER | 75.81% | 54.33% | 60.36% |
| CER | 54.99% | 42.41% | 46.18% | |
| TER | 54% | 17% | 27% | |
| Small | WER | 50.46% | 21.75% | 29.78% |
| CER | 35.89% | 16.58% | 22.37% | |
| TER | 43% | 8% | 18% |
Key Observations:
Note: To compute CER and WER, whitespaces are added between words in Khmer text (Khmer text does not have word boundaries like English text). To do so, khmercut PyPI package is used to tokenize Khmer text into words, and then the words are joined back together with whitespaces.
WER Comparison with Whisper:
| Tiny | Parameters | Khmer (fleurs) | English (librispeech.clean) |
|---|---|---|---|
| TrorYongASR | 29M | 75.88% | 54.33% |
| Whisper | 39M | 100.6% | 7.6% |
| Small | Parameters | Khmer (fleurs) | English (librispeech.clean) |
|---|---|---|---|
| TrorYongASR | 135M | 50.46% | 21.75% |
| Whisper | 244M | 104.4% | 3.4% |
Key Observations:
Note: WER data of Whisper is taken from their paper.
Language Detection: Both model sizes achieved great performance across all metrics (Precision, Recall, F1-score) on both datasets, indicating excellent binary classification capability for distinguishing between Khmer and English audio. This high score is expected because during pre-training, the model performs permutations on word tokens starting from position 3, while the first three positions (start token, language token, and task token) remain fixed. Since language detection relies on the language token at position 1, and this token is never permuted during pre-training, the model can achieve perfect accuracy on language detection tasks.
Transcription: The Small model shows strong performance on English (21.75% WER, 16.58% CER, 8% TER) and moderate performance for Khmer (50.46% WER, 35.89% CER, 43% TER). The Tiny model shows strong performance on English (54.33% WER, 42.41% CER, 17% TER) but significantly lower performance for Khmer (75.88% WER, 54.99% CER, 54% TER). This shows that TrorYongASR can be scaled to get higher performance.
Note on Translation Task: The models are also trained for translation task, but evaluation is deferred to future work due to scarce data (there are only 2000 examples from Khmer audio to English text, and 1000 examples from English audio to Khmer text in the pre-training).
First, install tror-yong-asr PyPI package:
pip install tror-yong-asr
Then, use the code below to get started with the model.
from transformers import AutoProcessor
from tror_yong_asr import TrorYongASRModel, transcribe, translate, detect_language
model_id = "KrorngAI/TrorYongASR-small"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = TrorYongASRModel.from_pretrained(model_id, trust_remote_code=True)
result1 = detect_language('/path/to/audio_file.mp3', model, processor)
print(result1)
result2 = transcribe('/path/to/audio_file.mp3', model, processor, max_tokens=64)
print(result2)
result3 = translate('/path/to/audio_file.mp3', model, processor, max_tokens=64)
print(result3)
Notebook (TBA)
The Tiny model can be used directly for:
The model can be integrated into:
Technical Limitations:
<|nospeech|> token.)Sociotechnical Limitations:
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
To capture model's scalability, both tiny and small variants were trained using the same configuration detailed below.
For transcription task, the model was trained on around 140 hours of Khmer audio and around 100 hours of English audio.
Khmer datasets include DDD-Cambodia/khm-asr-cultural (134.6 hours), openslr/openslr, and google/fleurs.
Split clean.100 of openslr/librispeech_asr was used for English dataset.
| Dataset | Language | Training examples | Validation examples | Description |
|---|---|---|---|---|
| DDD-Cambodia/khm-asr-cultural | Khmer | 56716 | 0 | Khmer ASR Cultural Dataset (split train) |
| openslr/openslr | Khmer | 2906 | 0 | Multi-speaker TTS data for Khmer language (split SLR42) |
| google/fleurs | Khmer | 1675 | 324 | TTS data for Khmer language (split km_kh) |
| librispeech_asr.clean | English | 28539 | 2703 | Clean speech dataset for English transcription |
For translation task, the data was scarce: only 2000 examples for Khmer audio to English text, and only 1000 examples for English audio to Khmer text.
Following Whisper model of openai, audios with duration longer than 30 seconds are filtered out.
All audios have 16000 sample rate.
For English dataset, all texts are in lowercase.
LightningAI package <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->The training was conducted over 3812 optimizer steps.
BibTeX:
@online{khun2026,
author = {Khun, Kimang},
title = {TrorYongASR: {Permuted} {AutoRegressive} {Sequence}
{Modeling} for {Automatic} {Speech} {Recognition}},
date = {2026-05-07},
url = {https://kimang18.github.io/krorngai-blog/TrorYongASR/},
langid = {en}
}
LightningAI, Kaggle, and Google Colab did not specifically sponsor this project.
But, both models are trained thanks to their free credits.
So, huge thanks to LightningAI, Kaggle and Google Colab.
Thanks to the authors of PARSeq and Whisper for their publicly available sourcecode.
Thanks to openslr, Mozilla Data Collective and Google for their publicly available dataset.
If you have any questions, please reach out at Facebook Page.