Downloads · 30 days
74
22% of all-time downloads
PavonicAI/HeartMuLa-3B-4bit
HeartMuLa-3B-4bit is a machine learning model from PavonicAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Pre-quantized 4-bit (NF4) checkpoint of HeartMuLa-oss-3B for 16 GB VRAM GPUs (RTX 4060 Ti, RTX 5070 Ti, etc.).
Downloads · 30 days
74
22% of all-time downloads
All-time downloads
344
Public
Parameters
4B
4.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors4.9 GB · 99%
How the weights are stored.
U83.2B · 79%
From the Hugging Face model README
Pre-quantized 4-bit (NF4) checkpoint of HeartMuLa-oss-3B for 16 GB VRAM GPUs (RTX 4060 Ti, RTX 5070 Ti, etc.).
All songs generated with this checkpoint on an RTX 5070 Ti (16 GB) using our ForgeAI ComfyUI Node:
| Song | Genre | Duration | CFG |
|---|---|---|---|
| Codigo del Alma (CFG 2) | Spanish Pop, Emotional | 3:00 | 2.0 |
| Codigo del Alma (CFG 3) | Spanish Pop, Emotional | 3:00 | 3.0 |
| Codigo del Alma (60s) | Spanish Pop | 1:00 | 2.0 |
| Codigo del Alma (Latin) | Latin Pop | 1:00 | 2.0 |
| Runtime | Chill, R&B | 3:00 | 2.0 |
| Forged in Code | Country Pop | 2:00 | 2.0 |
| Digital Rain | Electronic | 1:00 | 2.0 |
| Pixel Life | Pop | 1:00 | 2.0 |
The original HeartMuLa 3B model requires ~15 GB VRAM in bfloat16. Together with HeartCodec (~1.5 GB), it exceeds 16 GB VRAM, making it impossible to run on consumer GPUs like RTX 4060 Ti, RTX 5070 Ti, etc.
On top of that, the original code has several compatibility issues with modern PyTorch/transformers/torchtune versions (see fixes below).
Use our ForgeAI HeartMuLa ComfyUI Node for the easiest setup. All compatibility fixes are applied automatically.
Also available on the ComfyUI Registry.
Install via ComfyUI Manager or clone into custom_nodes:
cd ComfyUI/custom_nodes
git clone https://github.com/PavonicAI/ForgeAI-HeartMuLa.git
pip install -r ForgeAI-HeartMuLa/requirements.txt
Download this checkpoint into your ComfyUI models folder:
ComfyUI/models/HeartMuLa/HeartMuLa-oss-3B/
You still need the original HeartCodec and tokenizer from the original repo:
ComfyUI/models/HeartMuLa/
├── HeartMuLa-oss-3B/ ← this checkpoint
├── HeartCodec-oss/ ← from original repo
├── tokenizer.json ← from original repo
└── gen_config.json ← from original repo
HeartMuLa uses comma-separated tags to control style. Genre is the most important tag — always put it first.
genre:pop, emotional, synth, warm, female voice
| CFG | Best For | Notes |
|---|---|---|
| 2.0 | Pop, Ballads, Emotional | Sweet spot for clean vocals |
| 3.0 | Rock, Latin, Uptempo | More energy |
| 4.0+ | Electronic, Dance | May introduce artifacts |
[intro]
[verse]
Your lyrics here...
[chorus]
Chorus lyrics...
[outro]
If you want to use this checkpoint without ComfyUI, you need to apply several code fixes manually. See the sections below.
Add ignore_mismatched_sizes=True to ALL from_pretrained() calls:
HeartCodec.from_pretrained(..., ignore_mismatched_sizes=True)
HeartMuLa.from_pretrained(..., ignore_mismatched_sizes=True)
In modeling_heartmula.py, add RoPE init to setup_caches():
def setup_caches(self, ...):
# ... existing cache setup ...
for m in self.modules():
if hasattr(m, "rope_init"):
m.rope_init()
m.to(device)
Offload model to CPU before codec decode:
self.model.cpu()
torch.cuda.empty_cache()
wav = self.audio_codec.detokenize(frames)
Replace torchaudio with soundfile:
import soundfile as sf
sf.write(save_path, wav_np, 48000)
from transformers import BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_quant_type="nf4",
)
model = HeartMuLa.from_pretrained(
"PavonicAI/HeartMuLa-3B-4bit",
quantization_config=bnb_config,
device_map="cuda:0",
ignore_mismatched_sizes=True,
)
Apache-2.0 (same as original)