Downloads · 30 days
0
matbee/sam-audio-large-onnx
sam-audio-large-onnx is a machine learning model from matbee. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
ONNX-converted models for SAM-Audio (facebook/sam-audio-large) - Meta's Semantic Audio Modeling for audio source separation.
Downloads · 30 days
0
Access
Public
Updated Dec 23, 2025
Repo size
22.3 GB
Likes
10
Public
Click a slice to open those files.
.data21.9 GB · 100%
From the Hugging Face model README
ONNX-converted models for SAM-Audio (facebook/sam-audio-large) - Meta's Semantic Audio Modeling for audio source separation.
This repository contains both FP32 and FP16 versions of the models.
| Variant | DiT Size | Total Size | Notes |
|---|---|---|---|
fp32/ | 11.76 GB | ~13.9 GB | Full precision |
fp16/ | 5.88 GB | ~8.0 GB | Half precision (recommended) |
| File | Description | FP32 Size | FP16 Size |
|---|---|---|---|
dacvae_encoder.onnx | Audio encoder (48kHz → latent) | 110 MB | 110 MB |
dacvae_decoder.onnx | Audio decoder (latent → 48kHz) | 320 MB | 320 MB |
t5_encoder.onnx | Text encoder (T5-base) | 440 MB | 440 MB |
dit_single_step.onnx | DiT denoiser (3B params) | 11.76 GB | 5.88 GB |
vision_encoder.onnx | Vision encoder (CLIP-based) | 1.27 GB | 1.27 GB |
tokenizer/ | SentencePiece tokenizer files | - | - |
pip install onnxruntime sentencepiece torchaudio torchvision torchcodec soundfile
# For CUDA support (recommended for large model):
pip install onnxruntime-gpu
python onnx_inference.py \
--video input.mp4 \
--text "a person speaking" \
--model-dir fp16 \
--output target.wav \
--output-residual residual.wav
python onnx_inference.py \
--video input.mp4 \
--text "keyboard typing" \
--model-dir fp32 \
--output target.wav
python onnx_inference.py \
--audio input.wav \
--text "drums" \
--model-dir fp16 \
--output drums.wav
python -m onnx_export.export_dit \
--output-dir ./my_models \
--model-id facebook/sam-audio-large \
--fp16 \
--device cuda
python -m onnx_export.export_dacvae --output-dir ./my_models --model-id facebook/sam-audio-large
python -m onnx_export.export_t5 --output-dir ./my_models --model-id facebook/sam-audio-large
python -m onnx_export.export_vision --model facebook/sam-audio-large --output ./my_models
SAM-Audio is released under the CC-BY-NC 4.0 license. See original repository for full terms.
Original model by Meta AI Research.