Downloads · 30 days
0
BSC-LT/ConMamba-small-ca
ConMamba-small-ca is a automatic speech recognition model from BSC-LT. Use it when you need speech turned into text. The card lists the license as gpl-3.0.
<details <summaryClick to expand</summary
Downloads · 30 days
0
Access
Public
Updated Aug 20, 2026
Repo size
209 MB
Likes
0
Public
Click a slice to open those files.
.ckpt156 MB · 100%
From the Hugging Face model README
The ConMamba-small-ca is an acoustic model for Automatic Speech Recognition (ASR) in Catalan. It is based on the ConMamba architecture, which uses a Mamba (State Space Model) encoder augmented with convolutions for efficient sequence processing.
The ConMamba-small-ca model implements the Convolution-augmented Mamba (ConMamba) architecture, an adaptation of State Space Models (SSMs) designed to improve performance and efficiency in speech recognition tasks by integrating convolutional layers.
This model has been specifically trained for the Catalan language. The corpus used for training has 4929 hours.
This model can be used for Automatic Speech Recognition (ASR) in Catalan. The model is intended to transcribe audio files in Catalan to plain text without punctuation.
The implementation of the ConMamba-small-ca architecture often depends on specific libraries such as mamba-ssm and causal-conv1d. It is recommended to follow the installation steps from the original Mamba ASR repository:
conda create --name mamba_asr python=3.9
conda activate mamba_asr
clone github https://github.com/langtech-bsc/ConMamba_ASR
cd ConMamba_ASR
pip install -r requirements.txt
# Make sure that the versions of torch, torchaudio, causal-conv1d, and mamba-ssm are compatible with your hardware.
Inference is performed using the dedicated run_inference.py script provided within the repository.
# Define your paths
REPO_PATH="/path/to/ConMamba-ASR"
AUDIO="/path/to/your/audio.wav"
HPARAMS="conmambamamba_debug_catalan_small_1k_unigram_inference.yaml" # Use your specific inference YAML
# Execute inference script
python $REPO_PATH/run_inference.py \
--hparams $HPARAMS \
--audio $AUDIO
Dev Result - WER: 8.6
The model was trained for a total of 4929 hours. Including:
If this model contributes to your research, please cite the work:
@inproceedings{zevallos2025conmambasmallca,
title={Evaluating High-Performance and Lightweight ASR Systems for Catalan},
author={Zevallos, Rodolfo}
organization={Barcelona Supercomputing Center},
year={2025}
}
The model was trained during September (2025) in the Language Technologies Laboratory of the Barcelona Supercomputing Center by Rodolfo Zevallos.
For further information, please send an email to [email protected].
Copyright(c) 2025 by Language Technologies Laboratory, Barcelona Supercomputing Center.
This work has been promoted and financed by the Generalitat de Catalunya through the Aina project.
The conversion of the model was possible thanks to the computing time provided by Barcelona Supercomputing Center through MareNostrum 5.