Downloads · 30 days
8
36% of all-time downloads
voicing-ai/voicing-aligner
voicing-aligner is a machine learning model from voicing-ai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
--- license: apache-2.0 pipelinetag: automatic-speech-recognition
Downloads · 30 days
8
36% of all-time downloads
All-time downloads
22
Public
Parameters
918M
1.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.8 GB · 100%
From the Hugging Face model README
license: apache-2.0 pipeline_tag: automatic-speech-recognition
Voicing-aligner is a family of automatic speech recognition (ASR) models designed for accurate, fast, and robust speech recognition across 52 languages and dialects. The models support both language identification and automatic speech recognition, leveraging large-scale speech training data and the advanced audio understanding capabilities of the Voicing-Omni foundation model.
The 0.7B-parameter version delivers state-of-the-art performance among open-source ASR models and demonstrates competitive performance against leading proprietary commercial speech recognition APIs.
🌍 Multilingual Support Supports language identification and speech recognition across 52 languages and dialects, making it suitable for multilingual and global applications.
⚡ High Quality and Low Latency Provides accurate and robust transcription across a wide range of acoustic conditions, including noisy environments, diverse speakers, and challenging speech patterns.
🎯 Strong Recognition Performance Achieves competitive results across both open-source and internal evaluation benchmarks, with the 1.8B model delivering state-of-the-art performance among open-source ASR systems.
🚀 Production-Ready Inference Toolkit Alongside the model architectures and weights, Voicing-ASR provides a comprehensive inference framework designed for real-world deployment.
The toolkit supports:
Voicing-ASR is designed to provide a complete solution for researchers and developers looking to integrate high-quality multilingual speech recognition into production applications.