Downloads Ā· 30 days
0
Dzikriii/vocallet
vocallet is a machine learning model from Dzikriii. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A production-grade Python FastAPI microservice that processes voice commands for the Vocallet accessible finance application. Built for Indonesian users including visually impaired individuals, UMKM small businesses,ā¦
Downloads Ā· 30 days
0
Access
Public
Updated Jul 10, 2026
Repo size
ā
Likes
0
Public
Click a slice to open those files.
.md15.7 KB Ā· 56%
From the Hugging Face model README
A production-grade Python FastAPI microservice that processes voice commands for the Vocallet accessible finance application. Built for Indonesian users including visually impaired individuals, UMKM small businesses, and personal finance management.
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā CLIENT (React 19 / Node.js) ā
ā POST /api/voice-command (multipart audio) ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāā¬āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā
ā¼
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā VOCALLET AI SERVICE (FastAPI) ā
ā āāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāā ā
ā ā AudioServiceā ā STT Service ā ā Intent Service ā ā
ā ā validate āāāā¶ā (Whisper STT) āāāā¶ā (XLM-RoBERTa XNLI) ā ā
ā ā save temp ā ā transcribe() ā ā classify_intent() ā ā
ā ā load 16kHz ā ā ā ā zero-shot labels ā ā
ā āāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāā ā
ā ā
ā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā ā
ā ā ModelRegistry (Singleton) ā ā
ā ā _stt_pipeline (Whisper) ā _intent_pipeline (RoBERTa) ā ā
ā ā Loaded ONCE at startup via asyncio.gather() ā ā
ā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
| Requirement | Version | Notes |
|---|---|---|
| Python | 3.10+ | Required for match statements |
| pip | 23.0+ | For dependency resolution |
| ffmpeg | Any recent | Required for MP3/M4A/WebM processing |
| RAM | 4GB minimum | For Whisper Small + XLM-RoBERTa Large |
| RAM | 8GB recommended | For smooth concurrent inference |
| NVIDIA GPU | Optional | Enables faster inference; CUDA 11.8+ |
# Ubuntu/Debian
sudo apt-get install ffmpeg
# macOS (Homebrew)
brew install ffmpeg
# Windows (Chocolatey)
choco install ffmpeg
# Windows (Scoop)
scoop install ffmpeg
cd vocallet-ai-service
python -m venv venv
# Activate (Linux/macOS)
source venv/bin/activate
# Activate (Windows)
venv\Scripts\activate
# CPU-only (recommended for most users)
pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt
# GPU (CUDA 12.1) ā skip if using CPU
pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
cp .env.example .env
# Edit .env as needed (optional ā defaults work out of the box)
uvicorn app.main:app --host 0.0.0.0 --port 8001 --reload
The service will be available at:
http://localhost:8001/docshttp://localhost:8001/redochttp://localhost:8001/ui/index.htmlhttp://localhost:8001/healthā ļø First run note: On first startup, models (~1.3GB total) will download from Hugging Face. This can take 5ā20 minutes depending on your internet connection. The service will not accept requests until downloads complete.
| Variable | Default | Description |
|---|---|---|
APP_NAME | Vocallet AI Service | Service display name |
APP_VERSION | 1.0.0 | Service version string |
DEBUG | false | Enable debug mode + verbose logging |
HOST | 0.0.0.0 | Bind host address |
PORT | 8001 | Bind port |
ALLOWED_ORIGINS | http://localhost:3000,... | Comma-separated CORS allowed origins |
HF_TOKEN | (empty) | Hugging Face access token (optional for public models) |
HF_CACHE_DIR | ./models_cache | Local directory for model weight cache |
STT_MODEL_NAME | openai/whisper-small | Hugging Face model ID for speech-to-text |
INTENT_MODEL_NAME | joeddav/xlm-roberta-large-xnli | Hugging Face model ID for intent classification |
DEVICE | auto | Inference device: auto, cpu, cuda, cuda:0, mps |
TORCH_DTYPE | float32 | float32 (default) or float16 (GPU, faster + less VRAM) |
MAX_AUDIO_SIZE_MB | 25 | Maximum allowed upload file size in MB |
MAX_AUDIO_DURATION_SECONDS | 60 | Maximum allowed audio duration in seconds |
INTENT_CONFIDENCE_THRESHOLD | 0.45 | Scores below this fall back to "tidak diketahui" |
INTENT_LABELS | catat pengeluaran,... | Comma-separated intent label strings |
TEMP_DIR | ./temp_audio | Directory for temporary audio file storage |
POST /api/voice-commandProcess a voice audio file and return the transcript and intent.
Request:
Content-Type: multipart/form-data
Body: file=<audio_file>
curl -X POST http://localhost:8001/api/voice-command \
-F "file=@recording.wav;type=audio/wav"
Successful Response (200):
{
"status": "success",
"transcript": "tolong catat jualan basreng 50 ribu",
"transcript_language": "id",
"intent": "catat penjualan umkm",
"confidence_score": 0.9241,
"all_intent_scores": {
"catat penjualan umkm": 0.9241,
"catat pengeluaran": 0.0412,
"hitung zakat": 0.0214,
"baca laporan": 0.0091,
"tidak diketahui": 0.0042
},
"processing_time_ms": 1234.56,
"audio_duration_seconds": 3.2,
"model_info": {
"stt_model": "openai/whisper-small",
"intent_model": "joeddav/xlm-roberta-large-xnli",
"device": "cpu"
}
}
Error Response (4xx/5xx):
{
"status": "error",
"error_code": "AUDIO_TOO_LONG",
"message": "Audio duration (75.0s) exceeds maximum allowed (60s).",
"detail": "...",
"timestamp": "2024-08-10T12:34:56.789Z"
}
Error Codes:
| Code | HTTP | Description |
|---|---|---|
FILE_TOO_LARGE | 400 | File size exceeds MAX_AUDIO_SIZE_MB |
UNSUPPORTED_FORMAT | 400 | File format not in allowed list |
AUDIO_TOO_LONG | 400 | Audio duration exceeds MAX_AUDIO_DURATION_SECONDS |
SERVICE_UNAVAILABLE | 503 | Models not loaded yet |
TRANSCRIPTION_FAILED | 422 | Whisper could not transcribe the audio |
CLASSIFICATION_FAILED | 422 | Intent classification failed |
INTERNAL_ERROR | 500 | Unexpected server error |
GET /healthReturns service health and model loading status.
curl http://localhost:8001/health
{
"status": "healthy",
"service": "Vocallet AI Service",
"version": "1.0.0",
"uptime_seconds": 3600.5,
"models": {
"stt": {
"name": "openai/whisper-small",
"loaded": true,
"load_time_seconds": 12.3,
"device": "cpu",
"error": null
},
"intent": {
"name": "joeddav/xlm-roberta-large-xnli",
"loaded": true,
"load_time_seconds": 25.7,
"device": "cpu",
"error": null
}
}
}
Status values: healthy (both models loaded) | degraded (one model failed) | unhealthy (both failed)
GET /health/modelsDetailed model information including GPU memory usage (if applicable).
curl http://localhost:8001/health/models
GET /health/pingSimple liveness probe for Docker/Kubernetes healthchecks.
curl http://localhost:8001/health/ping
# Response: {"pong": true}
| Model | Size | Speed | Quality | Recommended Use |
|---|---|---|---|---|
whisper-tiny | ~75 MB | ā”ā”ā” Fastest | ā Basic | Development / testing |
whisper-base | ~145 MB | ā”ā” Fast | āā Decent | Low-resource deployment |
whisper-small | ~242 MB | ā” Good | āāā Good | Recommended ā |
whisper-medium | ~769 MB | š¢ Slow | āāāā High | High-accuracy use cases |
whisper-large-v3 | ~1.5 GB | š¢š¢ Slowest | āāāāā Best | Maximum accuracy |
| Model | Size | Accuracy | Languages |
|---|---|---|---|
joeddav/xlm-roberta-large-xnli | ~1.1 GB | āāāāā Best | 100+ (incl. Indonesian) ā |
cross-encoder/nli-MiniLM2-L6-H768 | ~120 MB | āāā Good | Primarily English |
| Format | Extension | MIME Type | Notes |
|---|---|---|---|
| WAV | .wav | audio/wav | Recommended, no conversion needed |
| MP3 | .mp3 | audio/mp3, audio/mpeg | Requires ffmpeg |
| OGG Vorbis | .ogg | audio/ogg | Web standard |
| WebM | .webm | audio/webm | Browser MediaRecorder default |
| FLAC | .flac | audio/flac | Lossless |
| M4A/AAC | .m4a | audio/m4a | iOS recording format |
Models are downloaded from Hugging Face Hub on first run:
whisper-small ā ~242MBxlm-roberta-large-xnli ā ~1.1GBAfter the first download, they are cached in ./models_cache/ and reused on subsequent restarts. To pre-warm the cache:
python -c "
from transformers import pipeline
pipeline('automatic-speech-recognition', 'openai/whisper-small', model_kwargs={'cache_dir': './models_cache'})
pipeline('zero-shot-classification', 'joeddav/xlm-roberta-large-xnli', model_kwargs={'cache_dir': './models_cache'})
"
Switch to CPU or a smaller model:
# In .env:
DEVICE=cpu
STT_MODEL_NAME=openai/whisper-tiny
Or use float16 for less VRAM usage (GPU only):
DEVICE=cuda
TORCH_DTYPE=float16
Ensure ffmpeg is installed and available in your PATH:
ffmpeg -version
If missing, install it:
# Ubuntu
sudo apt-get install ffmpeg
# macOS
brew install ffmpeg
Make sure you are inside the virtual environment:
source venv/bin/activate # Linux/macOS
venv\Scripts\activate # Windows
# Then reinstall:
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt
The audio may be:
Try recording a clear, audible voice command.
# Build the image
docker-compose build
# Start service
docker-compose up -d
# View logs
docker-compose logs -f vocallet-ai
# Stop service
docker-compose down
# Pass HF token securely via host environment
export HF_TOKEN=hf_your_token_here
docker-compose up -d
docker-compose ps
curl http://localhost:8001/health/ping
Example: Forwarding audio from Express to the Python microservice.
const axios = require('axios');
const FormData = require('form-data');
const multer = require('multer');
const upload = multer({ storage: multer.memoryStorage() });
// Express route: POST /voice-command
router.post('/voice-command', upload.single('audio'), async (req, res) => {
try {
const form = new FormData();
form.append('file', req.file.buffer, {
filename: req.file.originalname,
contentType: req.file.mimetype,
});
const response = await axios.post(
'http://localhost:8001/api/voice-command',
form,
{
headers: form.getHeaders(),
timeout: 60000 // 60s timeout for model inference
}
);
res.json(response.data);
} catch (err) {
const detail = err.response?.data?.detail || err.message;
res.status(err.response?.status || 500).json({
error: 'voice_processing_failed',
detail
});
}
});
# Install dev dependencies
pip install -r requirements-dev.txt
# Run all tests
pytest tests/ -v
# Run with coverage report
pytest tests/ -v --cov=app --cov-report=html
# Open coverage report
open htmlcov/index.html # macOS
xdg-open htmlcov/index.html # Linux
When you start the service for the first time, here is what happens:
detect_device() runs ā Auto-detects CUDA ā MPS ā CPUasyncio.gather() kicks off ā Both models begin downloading/loading concurrentlywhisper-small; goes to ./models_cache/xlm-roberta-large-xnli; goes to ./models_cache/GET /health returns "status": "healthy" ā Both models are liveā±ļø Total startup time on first run (CPU): 5ā25 minutes (mostly download) ā±ļø Total startup time on subsequent runs (cache warm): 30ā90 seconds (model load only)
| Method | Path | Description |
|---|---|---|
GET | / | Service info |
POST | /api/voice-command | Main endpoint ā process audio |
GET | /health | Readiness probe with model status |
GET | /health/models | Detailed model metadata |
GET | /health/ping | Liveness probe |
GET | /docs | Swagger UI |
GET | /redoc | ReDoc UI |
GET | /ui/index.html | Standalone HTML test UI |