Downloads · 30 days
0
justabigduck/NSFW_MMaudio
NSFW_MMaudio is a text-to-audio model from justabigduck. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains an FP16 safetensors version of the fine-tuned MMAudio model from cloud19/NSFWMMaudio, optimized for improved memory efficiency and faster loading times.
Downloads · 30 days
0
Access
Public
Updated Apr 12, 2026
Repo size
2.1 GB
Likes
1
Public
Click a slice to open those files.
.safetensors2.1 GB · 100%
From the Hugging Face model README
This repository contains an FP16 safetensors version of the fine-tuned MMAudio model from cloud19/NSFW_MMaudio, optimized for improved memory efficiency and faster loading times.
Base Model: cloud19/NSFW_MMaudio
Original Project: hkchengrex/MMAudio
large_44k (from the original MMAudio).safetensors)✅ ~50% smaller file size (FP32 → FP16 conversion)
✅ Faster loading with safetensors format
✅ Lower GPU memory usage during inference
✅ Same quality output (minimal precision loss with FP16)
✅ Better compatibility with modern ML frameworks
This model can be used as a drop-in replacement for the original model. Load the safetensors file instead of the original PyTorch checkpoint:
from safetensors.torch import load_file
# Load the FP16 model weights
model_weights = load_file("model_fp16.safetensors")
# Load into your MMAudio model architecture
# (follow the same usage pattern as the base model)
System Requirements:
For usage instructions, please refer to the base model repository and simply replace the model loading with the FP16 safetensors version.
Base Model: cloud19/NSFW_MMaudio
Original MMAudio: hkchengrex/MMAudio
Optimization: FP16 conversion for improved efficiency
All credit for the original architecture, fine-tuning, and model development goes to the respective authors. This repository only provides format optimization.
@inproceedings{cheng2025taming,
title={{MMAudio}: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis},
author={Cheng, Ho Kei and Ishii, Masato and Hayakawa, Akio and Shibuya, Takashi and Schwing, Alexander and Mitsufuji, Yuki},
booktitle={CVPR},
year={2025}
}