Downloads · 30 days
26
19% of all-time downloads
appautomaton/re-use-semamba-mlx
re-use-semamba-mlx is a audio-to-audio model from appautomaton. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
[](https://github.com/appautomaton/mlx-speech) [](https://appautomaton.renocrypt.com) [](https://huggingface.co/appautomaton/dramabox-tts-3.3b-bf16-mlx)
Downloads · 30 days
26
19% of all-time downloads
All-time downloads
136
Public
Parameters
9.6M
38.6 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors38.6 MB · 100%
From the Hugging Face model README
Pure-MLX conversion of NVIDIA RE-USE, a
~9.6M-parameter SEMamba universal speech-enhancement model. In
mlx-speech it cleans a voice
reference before VAE conditioning when DramaBox TTS
runs with denoise_ref=True, giving the cloning model a clean speaker anchor.
Non-commercial weights. These weights derive from NVIDIA RE-USE, licensed under the NVIDIA Source Code License (non-commercial). See the License section.
nvidia/RE-USE (SEMamba, bidirectional Mamba over STFT magnitude + phase)denoise_ref=True. Optional, off by default..safetensors (1416 keys, ~9.6M params). No quantization, no architecture change.mamba_ssm selective_scan_ref reference math, so no CUDA kernels (mamba-ssm / causal-conv1d) are required.| File | Component | Format | Size |
|---|---|---|---|
model.safetensors | SEMamba enhancer | fp32 | ~38 MB |
config.json | Model + STFT config | JSON | n/a |
Used automatically by DramaBox when you opt in:
import mlx_speech
tts = mlx_speech.tts.load("dramabox")
result = tts.generate(
"Voice cloning from a noisy reference.",
reference_audio="noisy_speaker.wav",
denoise_ref=True, # cleans the reference with this model first
)
tts.load("dramabox") resolves these weights automatically. To run the enhancer
directly:
hf download appautomaton/re-use-semamba-mlx --local-dir models/reuse/mlx
from pathlib import Path
from mlx_speech.generation.reuse import REUSEEnhancer
enhancer = REUSEEnhancer.from_dir(Path("models/reuse/mlx"))
clean = enhancer.enhance(noisy_waveform, in_sr=16000) # mono in, mono out
Denoising a short voice-reference clip before voice cloning, so the model conditions on a clean speaker/style anchor rather than the recording's noise. The enhancer runs on the reference input, never on generated output, so the TTS model's paralinguistic events (breaths, laughs) are preserved.
appautomaton/mlx-speechappautomaton/dramabox-tts-3.3b-bf16-mlxNVIDIA Source Code License (non-commercial). These weights are a format
conversion of nvidia/RE-USE and remain
governed by NVIDIA's license terms; by downloading or using them you agree to
those terms. They may not be used commercially. Set denoise_ref=False (the
default) to run DramaBox voice cloning without this model. The mlx-speech
runtime code is MIT.