Downloads · 30 days
17
39% of all-time downloads
FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1
FastVideo-Minimax-FastH3-Preview-v0.1 is a text-to-video model from FastVideo. Use it when you need video from a text prompt. It is set up for diffusers. The card lists the license as other.
A few-step (4-step) distillation preview of MiniMax-H3, the 33B dual-modality (video + audio) diffusion transformer — distilled with data-free DMD2 by the FastVideo team.
Downloads · 30 days
17
39% of all-time downloads
All-time downloads
44
Public
Parameters
35B
148 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors148 GB · 100%
From the Hugging Face model README
A few-step (4-step) distillation preview of MiniMax-H3, the 33B dual-modality (video + audio) diffusion transformer — distilled with data-free DMD2 by the FastVideo team.
The base model samples with 50 denoising steps; this student walks a 4-step grid on the release's shift-12 rectified-flow schedule (12.5× fewer transformer evaluations), generating synchronized video and audio in one pipeline call.
Preview status (v0.1): this is an early training checkpoint (step 1400 of a 4000-step run) published for evaluation and integration work. Sample quality is still maturing; expect a stronger release checkpoint from the same run.
Diffusers-format (modular pipeline) layout. Only the transformer/ weights differ
from the base release — the distilled student, in bf16. All other components
(Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are
unmodified copies of the base release, included so the repo is self-contained.
The student was trained with block-sparse video attention (VSA, 64-token tiles,
90% sparsity) and carries its trained sparse-gate parameters
(attn.to_gate_compress); it can be run dense (default) or with VSA for
additional inference speedup.
from fastvideo import VideoGenerator
gen = VideoGenerator.from_pretrained(
"FastVideo/FastVideo-Minimax-FastVideo-Minimax-FastH3-Preview-v0.1",
num_gpus=1,
)
video = gen.generate_video(
prompt="<your H3-format multimodal prompt>",
num_inference_steps=4, # the distilled grid
guidance_scale=1.0, # the base model is guidance-distilled
)
Prompts follow the MiniMax-H3 multimodal prompt format
(integrated_multimodal_description: ... overall_soundscape: ...); see the base
model card for the prompting guide.
Distributed under the MiniMax H3 Community License (see LICENSE), inherited from the base model. Review the license (including its territory and acceptable-use terms) before use or redistribution.
transformer_ref component (reference-conditioning variant) is not
packaged here; its entry in modular_model_index.json points at the base
MiniMaxAI/MiniMax-H3 repo and is fetched from there if used. This preview
distills the text-to-video+audio path only.We thank Nuva Lab for bringing production grounding to FastH3 through its experience with real-world creative video-agent workloads. Its production-aligned post-training insights help bridge open-source research to practical data-assisted distillation for commercial video workflows, with Omni Ref as the next focus.
We thank the NVIDIA FastGen team for the DMD2 framework and H3 reference experiment that helped us align the score clock, modality shifts, and backward simulation.
We also thank MiniMax for releasing H3-Base, and the vLLM project, NVIDIA, and MBZUAI for their continued sponsorship and support of FastVideo.