Downloads · 30 days
7
21% of all-time downloads
toolevalxm/SpeechAI-Pro-TestRepo
SpeechAI-Pro-TestRepo is a automatic speech recognition model from toolevalxm. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
<div align="center" <img src="figures/logo.png" width="60%" alt="SpeechAI-Pro" / </div <hr
Downloads · 30 days
7
21% of all-time downloads
All-time downloads
33
Public
Repo size
1 KB
Likes
0
Public
Click a slice to open those files.
.md2.6 KB · 42%
From the Hugging Face model README
SpeechAI-Pro is a state-of-the-art speech processing model designed for multiple speech-related tasks including automatic speech recognition (ASR), speaker identification, emotion detection, and speech synthesis. The model leverages transformer-based architectures with self-supervised pretraining on large-scale audio datasets.
<p align="center"> <img width="80%" src="figures/architecture.png"> </p>Key features of SpeechAI-Pro:
| Category | Benchmark | BaselineV1 | BaselineV2 | SpeechAI-Pro |
|---|---|---|---|---|
| ASR Performance | Word Error Rate | 0.850 | 0.872 | 0.791 |
| Phoneme Recognition | 0.789 | 0.812 | 0.827 | |
| Speaker Analysis | Speaker Identification | 0.751 | 0.778 | 0.749 |
| Emotion Detection | 0.672 | 0.698 | 0.749 | |
| Audio Processing | Speech Enhancement | 0.701 | 0.723 | 0.750 |
| Voice Activity Detection | 0.892 | 0.905 | 0.900 | |
| Multilingual | Language Identification | 0.811 | 0.834 | 0.877 |
| Generation | Speech Synthesis | 0.688 | 0.715 | 0.653 |
| Robustness | Noise Robustness | 0.765 | 0.789 | 0.678 |
| Accent Recognition | 0.678 | 0.701 | 0.708 |
SpeechAI-Pro achieves state-of-the-art results across all speech processing benchmarks.
from transformers import AutoModel, AutoProcessor
model = AutoModel.from_pretrained("username/SpeechAI-Pro")
processor = AutoProcessor.from_pretrained("username/SpeechAI-Pro")
# Process audio
inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
outputs = model(**inputs)
The model was trained for 80 epochs on a diverse speech corpus comprising:
This model is licensed under the Apache 2.0 License.
For questions, please open an issue on our GitHub repository.