Downloads · 30 days
0
atharva-again/indic-conformer-600m-quantized
indic-conformer-600m-quantized is a automatic speech recognition model from atharva-again. Use it when you need speech turned into text. The card lists the license as mit.
This repository contains a quantized version of the Indic Conformer model, a large-scale automatic speech recognition (ASR) model created for Indic languages by AI4Bharat. The original model can be found here
Downloads · 30 days
0
Access
Public
Updated Dec 7, 2025
Repo size
676 MB
Likes
7
Public
Click a slice to open those files.
.onnx676 MB · 100%
From the Hugging Face model README
This repository contains a quantized version of the Indic Conformer model, a large-scale automatic speech recognition (ASR) model created for Indic languages by AI4Bharat. The original model can be found here
These benchmarks were conducted on Google Colab free tier with Tesla T4 GPU for Hindi.
You can use the notebooks in scripts directory to reproduce the results or compute for other languages.
| Decoding Method | FP 32 WER | int8 WER | FP32 CER | int8 CER |
|---|---|---|---|---|
| CTC | 0.1645 | 0.2985 | 0.0661 | 0.1698 |
| RNNT | 0.1508 | 0.2939 | 0.0642 | 0.149 |
This model is intended for transcribing speech in Indic languages into text. It can be used for applications such as voice assistants, transcription services, and accessibility tools.
To use this model, simply install the helper package:
pip install indic-asr-onnx
from indic_asr_onnx import IndicTranscriber
# Initialize (downloads model automatically)
transcriber = IndicTranscriber()
# Transcribe audio using CTC head
text = transcriber.transcribe_ctc("audio.wav", "hi") # Hindi
print(text)
# Transcribe audio using RNNT head
text = transcriber.transcribe_rnnt("audio.wav", "hi") # Hindi
print(text)
vocab.json: Subword vocabulary for supported languageslanguage_masks.json: Language-specific masks for handling multilingual inputsctc_decoder_quantized_int8.onnx: Quantized CTC decoder for connectionist temporal classificationencoder_quantized_int8.onnx: Quantized Conformer encoder for feature extraction from audiojoint_enc_quantized_int8.onnx: Quantized joint encoder component for RNN-T decodingjoint_pre_net_quantized_int8.onnx: Quantized joint pre-net for preprocessing in RNN-Tjoint_pred_quantized_int8.onnx: Quantized joint predictor for RNN-T decodingrnnt_decoder_quantized_int8.onnx: Quantized RNN-T decoder for recurrent neural network transduceradapters/*: Language-specific quantized joint post-net adapters for each supported language (e.g., joint_post_net_hi_quantized_int8.onnx for Hindi)Calibration Dataset:https://www.kaggle.com/datasets/haposeiz/indicvoices-calibration-1408
The Calibration Dataset was curated from the Indic Voices Dataset.
GitHub: https://github.com/atharva-again/indic-asr-onnx
For questions or issues, you can either open an issue on this repository, on GitHub, or email me at [email protected].