Skip to content

humanify

LongCat-AudioDiT-Env-TTS-1B-augment

humanify/LongCat-AudioDiT-Env-TTS-1B-augment

LongCat-AudioDiT-Env-TTS-1B-augment is a text-to-speech model from humanify. Use it when you need text read aloud. It is set up for transformers. The card lists the license as other.

Fine-tune of meituan-longcat/LongCat-AudioDiT-1B for the three-stream env-tts task: given a reference environment audio, a reference speaker audio, and three text streams (env caption / speaker caption / target speech…

Downloads · 30 days

4

5% of all-time downloads

All-time downloads

79

Public

Parameters

1.4B

5.7 GB on disk

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors5.7 GB · 100%

At a glance

Task
Text-to-Speech
Library
transformers
License
other
Model type
audiodit
Access
Public
Created
Jun 9, 2026
Updated
Jun 9, 2026
SHA
66d0e2c1

Base models

Task
Text-to-Speech
Library
transformers
Type
audiodit
License
other
Created
Jun 9, 2026
Updated
Jun 9, 2026