Downloads · 30 days
11
3% of all-time downloads
Vrspi/SpeechToText
SpeechToText is a automatic speech recognition model from Vrspi. Use it when you need speech turned into text. It is set up for transformers.
This model is designed to transcribe speech in the Moroccan dialect to text. It's built on top of the Wav2Vec 2.0 architecture, fine-tuned on a dataset of Moroccan dialect speech.
Downloads · 30 days
11
3% of all-time downloads
All-time downloads
430
Public
Parameters
315M
1.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
This model is designed to transcribe speech in the Moroccan dialect to text. It's built on top of the Wav2Vec 2.0 architecture, fine-tuned on a dataset of Moroccan dialect speech.
This model is part of a project aimed at improving speech recognition technology for underrepresented languages, with a focus on the Moroccan Arabic dialect. The model leverages the power of the Wav2Vec2 architecture, fine-tuned on a curated dataset of Moroccan speech.
This model is intended for direct use in applications requiring speech-to-text capabilities for the Moroccan dialect. It can be integrated into services like voice-controlled assistants, dictation software, or for generating subtitles in real-time.
This model is not intended for use with languages other than Moroccan Arabic or for non-speech audio transcription. Performance may significantly decrease when used out of context.
The model may exhibit biases present in the training data. It's important to note that dialectal variations within Morocco could affect transcription accuracy. Users should be aware of these limitations and consider additional validation for critical applications.
Continual monitoring and updating of the model with more diverse datasets can help mitigate biases and improve performance across different dialects and speaking styles.
from transformers import Wav2Vec2Processor, Wav2Vec2ForCTC
from transformers import pipeline
import soundfile as sf
# Load the model and processor
processor = Wav2Vec2Processor.from_pretrained("Vrspi/SpeechToText")
model = Wav2Vec2ForCTC.from_pretrained("Vrspi/SpeechToText")
# Create a speech-to-text pipeline
speech_recognizer = pipeline("automatic-speech-recognition", model=model, processor=processor)
# Load an audio file
speech, sampling_rate = sf.read("path_to_your_audio_file.wav")
# Transcribe the speech
transcription = speech_recognizer(speech, sampling_rate=sampling_rate)
print(transcription)
The model was trained on a dataset comprising approximately 20 hours of spoken Moroccan Arabic collected from various sources, including public speeches, conversations, and media content.
The audio files were resampled to 16kHz and trimmed to remove silence. Noisy segments were manually annotated and excluded from training.
The model is not tested yet , I will drop results as soon as possible