Downloads · 30 days
0
ConcertIDC/WhisperAI-Speech-To-Text
WhisperAI-Speech-To-Text is a machine learning model from ConcertIDC. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
- Multilingual Transcription: Automatically transcribes audio in various languages using OpenAI’s Whisper model. - Speaker Diarization: Detects different speakers in the audio and labels the transcription accordingly.…
Downloads · 30 days
0
Access
Public
Updated Jan 29, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py6.1 KB · 59%
From the Hugging Face model README
Make sure to have the following Python libraries installed. You can install them using pip and the requirements.txt file provided.
Clone the repository:
git clone https://github.com/your-repository-url.git
cd your-repository
Create and activate a virtual environment (optional but recommended):
python -m venv env
source env/bin/activate # On Windows, use `env\Scripts\activate`
Install the dependencies:
pip install -r requirements.txt
Install Hugging Face authentication token for pyannote audio (if required):
use_auth_token="your_token"
Run the Streamlit app:
streamlit run app.py
The app will launch in your browser. Select an audio file (MP3, WAV, or M4A format) from your system.
The file will be uploaded to the upload directory, and the transcription will begin.
After processing, the app will display:
whisper Python package.pyannote.audio library.If you face issues with loading the diarization model, ensure you have:
If the models fail to load, ensure that:
If the file upload is not working correctly:
upload folder exists in your project directory.