Downloads · 30 days
10
3% of all-time downloads
ModelsLab/midashenglm-gen-wer-lora
midashenglm-gen-wer-lora is a text-to-audio model from ModelsLab. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
A LoRA adapter for mispeech/midashenglm-gen, trained across all five of the model's capabilities — speech, sound effects, music, ambience and mixed scenes — rather than trading one for another.
Downloads · 30 days
10
3% of all-time downloads
All-time downloads
314
Public
Repo size
5.9 GB
Likes
0
Public
Click a slice to open those files.
.safetensors292 MB · 90%
From the Hugging Face model README
A LoRA adapter for mispeech/midashenglm-gen,
trained across all five of the model's capabilities — speech, sound effects,
music, ambience and mixed scenes — rather than trading one for another.
Same prompt, same seed in both arms — generate() seeds the global RNG, so the
solver draws identical noise and every audible difference is the adapter.
Speech, 4-word line — "The river remembers everything." The bucket that went 13.9% → 0.0% WER; the base arm garbles it.
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-4w__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-4w__V8.wav"></audio> |
Speech, 16-word line — both arms intelligible; the adapter is cleaner and tighter.
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-16w__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-16w__V8.wav"></audio> |
SFX — footsteps on gravel. Texture and caption match improve; the adapter clip runs shorter, the one real remaining cost.
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/sfx-gravel__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/sfx-gravel__V8.wav"></audio> |
Music — slow solo piano. The adapter plays longer and closer to the caption at this seed.
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/music-piano__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/music-piano__V8.wav"></audio> |
Ambience — steady rain on a metal roof. Near-identical by design: this arm held.
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/ambience-rain__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/ambience-rain__V8.wav"></audio> |
Two weak spots in the base model, measured rather than assumed.
Short lines. A four-word line scores around 42% WER where a sixteen-word one scores near zero. The failures are onsets: the model needs a moment to settle and a short line does not give it one. Prompt-side remedies make it worse, so it is a property of the weights.
Sound effects. The model's worst benchmark — AudioCaps FAD 5.01 against TangoFlux's 2.26 — which the paper attributes to training on mixed scenes rather than dedicated sound-effect corpora.
Same prompt, same seed, base weights against these adapters. generate() seeds
the global RNG, so the flow-matching solver draws identical noise in both arms
and every difference is the adapters.
| capability | metric | base | tuned | change |
|---|---|---|---|---|
| speech | WER | 5.6% | 0.2% | -5.4% |
| sfx | CLAP text | 0.4064 | 0.3837 | -0.0226 |
| sfx | CLAP real | 0.4152 | 0.4201 | +0.0049 |
| music | CLAP text | 0.4736 | 0.4673 | -0.0063 |
| music | CLAP real | 0.5928 | 0.6012 | +0.0084 |
| ambience | CLAP text | 0.2442 | 0.2070 | -0.0371 |
| ambience | CLAP real | 0.3667 | 0.3732 | +0.0065 |
| mixed | WER | 4.8% | 4.7% | -0.1% |
| mixed | CLAP text | 0.1708 | 0.1745 | +0.0037 |
| mixed | CLAP real | 0.7014 | 0.6943 | -0.0072 |
| speech | WER, 4-word lines | 13.9% | 0.0% | -13.9% |
| speech | WER, 9-word lines | 3.2% | 0.0% | -3.2% |
| speech | WER, 16-word lines | 1.9% | 0.6% | -1.2% |
| speech | WER, 27-word lines | 3.4% | 0.3% | -3.1% |
| speech | mean length | 4.73s | 4.47s | -0.26s |
| sfx | mean length | 8.08s | 6.93s | -1.15s |
| music | mean length | 9.42s | 9.19s | -0.23s |
| ambience | mean length | 10.43s | 10.29s | -0.14s |
| mixed | mean length | 4.97s | 4.47s | -0.50s |
CLAP text is "does it match the caption". CLAP real is cosine to the centroid of genuine MECAT audio of that category — "does it sound like the real thing", which text similarity alone cannot see. WER cannot see either: a clip with no words in it scores 0% however badly the room tone came out.
WER improved by 5.4 points, and that part is solid. The scene arms are not.
Caption adherence or realism fell on: sfx clap_text -0.023, ambience clap_text -0.037.
Use this for speech. For beds, effects and mixed scenes, A/B it against the base model on your own prompts first — the held-out flow loss improved on every capability, and on the scene arms that did not translate.
| Method | LoRA r=32, alpha=64, on LLM attention/MLP and the flow-matching DiT |
| Trainable | 72876032 of 2.89 B |
| Precision | float32 with TF32 matmuls, matching how the base model is served |
| Optimiser | AdamW-8bit, lr 5e-05, cosine decay, grad clip 1.0 |
| Steps | 2500 at an effective batch of 16 |
| Capability mix | sfx 22%, ambience 20%, speech 20%, mixed 20%, music 18% |
| Speech data | LibriTTS-R + MECAT S00; clips under 3 s oversampled 2.5x |
| Scene data | MECAT-Caption, all eight categories |
| Anchor | vector field pulled towards the frozen base, weight 0.5 on 35% of batches |
Base-model hyperparameters follow arXiv:2608.11804 §3.3 where they apply (lambda_stop = 0.01, grad clip 1.0).
Four things, because one is not enough:
from peft import PeftModel
from transformers import AutoModel
import soundfile as sf
model = AutoModel.from_pretrained("mispeech/midashenglm-gen", trust_remote_code=True)
model = PeftModel.from_pretrained(model, "ModelsLab/midashenglm-gen-wer-lora").merge_and_unload().cuda().eval()
result = model.generate(
"<|caption|> A close, clean recording of a single narrator, with nothing behind it. "
"<|asr|> The river remembers everything. "
"<|speech|> A woman in her forties, warm and even, speaking clearly. "
"<|sfx|> <|unknown|> <|music|> <|unknown|> <|env|> <|unknown|>",
seed=11,
)
sf.write("out.wav", result["audio"], result["sample_rate"])
Tag order matters more than anything else in the prompt: <|asr|> comes
before <|speech|>. Reversed, the model produces fluent unrelated speech —
14.2% mean WER against 373%.
Apache 2.0, following the base model. Training data: LibriTTS-R (CC-BY-4.0) and MECAT-Caption (CC-BY-3.0), both attribution-only.
Training code: audio-scenegen/training.