Downloads · 30 days
0
mistral-hackaton-2026/meetingmind-gpu
meetingmind-gpu is a audio classification model from mistral-hackaton-2026. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for custom.
GPU-accelerated speaker diarization and embedding extraction for the MeetingMind pipeline. Runs as an HF Inference Endpoint on a T4 GPU with scale-to-zero.
Downloads · 30 days
0
Access
Public
Updated Mar 1, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.py9.1 KB · 69%
From the Hugging Face model README
GPU-accelerated speaker diarization and embedding extraction for the MeetingMind pipeline. Runs as an HF Inference Endpoint on a T4 GPU with scale-to-zero.
GET /healthReturns service status and GPU availability.
curl -H "Authorization: Bearer $HF_TOKEN" $ENDPOINT_URL/health
{"status": "ok", "gpu_available": true}
POST /diarizeSpeaker diarization using pyannote v4. Accepts any audio format (FLAC, WAV, MP3, etc.).
curl -X POST \
-H "Authorization: Bearer $HF_TOKEN" \
-F [email protected] \
-F min_speakers=2 \
-F max_speakers=6 \
$ENDPOINT_URL/diarize
{
"segments": [
{"speaker": "SPEAKER_00", "start": 0.5, "end": 3.2, "duration": 2.7},
{"speaker": "SPEAKER_01", "start": 3.4, "end": 7.1, "duration": 3.7}
]
}
POST /embedSpeaker embedding extraction using FunASR CAM++. Returns L2-normalized 192-dim vectors for voiceprint matching.
curl -X POST \
-H "Authorization: Bearer $HF_TOKEN" \
-F [email protected] \
-F start_time=1.0 \
-F end_time=5.0 \
$ENDPOINT_URL/embed
{"embedding": [0.012, -0.034, ...], "dim": 192}
| Variable | Default | Description |
|---|---|---|
HF_TOKEN | (required) | Hugging Face token for pyannote model access |
PYANNOTE_MIN_SPEAKERS | 1 | Minimum speakers for diarization |
PYANNOTE_MAX_SPEAKERS | 10 | Maximum speakers for diarization |
pytorch/pytorch:2.4.0-cuda12.4-cudnn9-runtime