Downloads · 30 days
0
Panhapich/pre-trained_whisper_wavLM
pre-trained_whisper_wavLM is a automatic speech recognition model from Panhapich. Use it when you need speech turned into text. The card lists the license as mit.
This repository contains pretrained checkpoints and benchmarking results used to identify the most suitable speech encoder backbone for a Khmer Automatic Speech Recognition (ASR) system.
Downloads · 30 days
0
Access
Public
Updated Jul 3, 2026
Repo size
838 MB
Likes
1
Public
Click a slice to open those files.
.pt838 MB · 100%
From the Hugging Face model README
This repository contains pretrained checkpoints and benchmarking results used to identify the most suitable speech encoder backbone for a Khmer Automatic Speech Recognition (ASR) system.
The goal of this project was to compare the transferability of different self-supervised speech representations for Khmer ASR before investing in large-scale fine-tuning.
Three widely used pretrained speech encoders were evaluated:
openai/whisper-base)To ensure a fair comparison, each encoder was:
The resulting validation CTC loss was used as the primary evaluation metric.
For each backbone:
This approach evaluates how useful the pretrained speech representations are for Khmer ASR without any encoder fine-tuning.
| Encoder | Validation CTC Loss |
|---|---|
| Whisper Base | 0.5663 |
| WavLM Base | 0.6031 |
| Wav2Vec2 Base | 0.7836 |
🥇 Whisper Base — 0.5663
🥈 WavLM Base — 0.6031
🥉 Wav2Vec2 Base — 0.7836
Based on these experiments, Whisper Base produced the strongest transferable speech representations for Khmer ASR and was selected as the backbone for subsequent training stages.
{
"best_backbone_key": "whisper-base",
"best_backbone_id": "openai/whisper-base",
"hidden_size": 512,
"checkpoint_path": "./whisper-base_best.pt"
}
whisper-base_best.pt
The benchmark used a combination of multiple Khmer speech datasets.
Dataset:
seanghay/khmer_grkpp_speech
Dataset:
seanghay/km-speech-corpus
Dataset:
KrorngAI/fleurs_openslr42_mpwt
For balanced experimentation, up to 3,000 samples per dataset were used during backbone evaluation.
The model uses a character-level Khmer vocabulary containing:
[BLANK]
[PAD]
[MASK]
[UNK]
108 tokens
.
├── whisper-base_best.pt
├── README.md
└── benchmark metadata
This repository is not intended to be a production-ready ASR system.
Instead, it provides:
Planned improvements include:
If you use this repository in your research, please cite:
@misc{uk2026khmerasrbenchmark,
title={Khmer ASR Encoder Benchmark and Pretrained Checkpoints},
author={Uk, Panhapich},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/Panhapich/pre-trained_whisper_wavLM}
}
Panhapich Uk
Independent research project focused on:
This project is released under the MIT License.