Downloads · 30 days
0
0% of all-time downloads
marcosremar2/salmonn-inference
salmonn-inference is a machine learning model from marcosremar2. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A ready-to-use inference server for SALMONN (Speech Audio Language Music Open Neural Network) - a multimodal LLM that can understand speech, audio events, and music.
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
13
Public
Repo size
146 KB
Likes
0
Public
Click a slice to open those files.
.py268 KB · 61%
From the Hugging Face model README
A ready-to-use inference server for SALMONN (Speech Audio Language Music Open Neural Network) - a multimodal LLM that can understand speech, audio events, and music.
git clone https://huggingface.co/marcosremar2/salmonn-inference
cd salmonn-inference
# Create virtual environment
python -m venv venv
source venv/bin/activate
# Install dependencies
pip install -r requirements.txt
./install.sh
This downloads (~20GB):
python server.py
Server runs at http://localhost:8000
curl -X POST "http://localhost:8000/transcribe" \
-F "audio=@your_audio.wav"
curl -X POST "http://localhost:8000/chat" \
-F "audio=@your_audio.wav" \
-F "question=What is being said in this audio?"
from inference import SALMONNInference
model = SALMONNInference()
model.load()
# Transcribe
text = model.transcribe("audio.wav")
# Ask questions
answer = model.chat("audio.wav", "What language is being spoken?")
# Describe audio
description = model.describe("audio.wav")
| Endpoint | Method | Description |
|---|---|---|
/ | GET | API info |
/health | GET | Health check |
/transcribe | POST | Transcribe audio to text |
/chat | POST | Ask questions about audio |
/describe | POST | Get audio description |
Edit config.yaml to customize:
model:
device: "cuda:0" # GPU device
server:
host: "0.0.0.0"
port: 8000
generation:
max_new_tokens: 200
temperature: 1.0
Tested on NVIDIA L4 (24GB):
| Metric | Value |
|---|---|
| Model Load Time | ~20s |
| Audio Encode | ~250ms |
| Time to First Token | ~150ms |
| Tokens/second | ~18 |
| GPU Memory | ~16GB |
This repository uses Vicuna 7B v1.5 (not v1.1). The original SALMONN checkpoint was trained with v1.5, and using v1.1 will result in broken outputs (<unk> tokens).