Downloads · 30 days
44
15% of all-time downloads
Or4kool/wolof_asr
wolof_asr is a automatic speech recognition model from Or4kool. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
Fine-tuned wav2vec2 CTC model for transcribing Wolof speech to text.
Downloads · 30 days
44
15% of all-time downloads
All-time downloads
286
Public
Parameters
316M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
Fine-tuned wav2vec2 CTC model for transcribing Wolof speech to text.
AutoModelForCTC)wo)| Field | Requirement |
|---|---|
| Audio format | WAV, FLAC, OGG, MP3, M4A (decoded server-side) |
| Channels | Mono preferred (stereo is averaged to mono) |
| Sample rate | Any (resampled to 16 kHz automatically) |
| Duration | Up to 60 seconds per request |
| Encoding | Raw bytes, base64, multipart upload, or HTTPS URL |
from transformers import AutoModelForCTC, AutoProcessor
import librosa
model_id = "Or4kool/wolof_asr"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForCTC.from_pretrained(model_id)
model.eval()
audio, sr = librosa.load("sample.wav", sr=16000, mono=True)
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
logits = model(inputs.input_values).logits
predicted_ids = torch.argmax(logits, dim=-1)
text = processor.batch_decode(predicted_ids)[0]
print(text)
from transformers import pipeline
asr = pipeline(
"automatic-speech-recognition",
model="Or4kool/wolof_asr",
)
print(asr("sample.wav")["text"])
This model powers the myagro-audio service. Example requests:
curl -X POST https://your-endpoint/transcribe \
-H "Content-Type: application/json" \
-d '{
"audio_base64": "<base64-audio>",
"language": "wol",
"return_timestamps": false
}'
curl -X POST https://your-endpoint/transcribe/upload \
-F "[email protected]" \
-F "language=wol" \
-F "return_timestamps=true"
{
"input": {
"items": [
{"id": "clip-1", "audio_base64": "..."},
{"id": "clip-2", "audio_base64": "..."}
],
"return_timestamps": false
}
}
{
"text": "transcribed wolof text",
"language": "wol",
"duration_seconds": 2.5,
"request_id": "uuid",
"inference_ms": 120.4,
"words": [
{"word": "example", "start": 0.12, "end": 0.48}
]
}
words is optional and only included when return_timestamps=true.
@misc{or4kool-wolof-asr,
author = {Samuel Oriaku},
title = {Wolof ASR wav2vec2 CTC},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Or4kool/wolof_asr}}
}
huggingface-cli login
huggingface-cli upload Or4kool/wolof_asr huggingface_model_card/README.md README.md