Downloads · 30 days
3
10% of all-time downloads
smithcooly2k/Motif-Video-2B
Motif-Video-2B is a text-to-video model from smithcooly2k. Use it when you need video from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
<p align="center" <img src="assets/banner.png" width="100%" alt="Motif-Video 2B teaser"/ </p
Downloads · 30 days
3
10% of all-time downloads
All-time downloads
29
Public
Parameters
2B
41.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17 GB · 99%
From the Hugging Face model README
Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. Motif-Video 2B asks whether competitive text-to-video quality is reachable at a much smaller budget — fewer than 10M training clips and under 100,000 H200 GPU hours — and shows that the answer is yes, provided the model design explicitly separates objectives that scaling would otherwise leave entangled.
Our central observation is that prompt alignment, temporal consistency, and fine-detail recovery interfere with one another when handled through the same pathway. Motif-Video 2B addresses this objective interference architecturally rather than relying on scale alone, through two contributions:
These are paired with a micro-budget training recipe combining TREAD token routing and early-phase REPA with a frozen V-JEPA teacher — to our knowledge, the first time this combination has been applied to text-to-video training.
On VBench, Motif-Video 2B reaches 83.76%, the highest Total Score among open-source models we evaluate, surpassing Wan2.1-14B at 7× fewer parameters and roughly an order of magnitude less training data.
<!-- Architecture figure — replace with Figure 2 from the technical report (the three-stage backbone + Shared Cross-Attention diagram). --> <p align="center"> <img src="assets/architecture.png" width="90%" alt="Motif-Video 2B architecture"/> </p>Motif-Video 2B is a flow-matching diffusion transformer organized around a single principle: each component is assigned a well-defined responsibility, and components with conflicting objectives are not asked to share capacity.
| Component | Choice |
|---|---|
| Text encoder | T5Gemma2 (encoder–decoder, UL2-adapted Gemma 3) |
| Video tokenizer | Wan2.1 VAE (8×8 spatial, 4× temporal compression), 2×2×1 patchify |
| Backbone | 12 dual-stream + 16 single-stream + 8 DDT decoder layers |
| Hidden dim / heads | 1536 / 12 heads × 128 |
| Normalization | QK-normalization throughout |
| Position encoding | RoPE |
| Cross-attention | Shared Cross-Attention in the single-stream stage |
| Objective | Rectified flow matching (velocity prediction) |
| I2V conditioning | First-frame latent + SigLIP image embeddings, with timestep-aware blur |
A high-level walkthrough of the role separation:
For the full derivation of why Shared Cross-Attention shares K/V but not Q, and why this is necessary in addition to standard zero-init of W_O, see Section 3.3 of the technical report.
<!-- Optional: insert Figure 3 (attention heatmaps across the three stages) here as a secondary architecture figure. It is the strongest visual evidence for the role-separation argument. -->pip install "transformers>=5.5.4" torch accelerate ftfy einops sentencepiece regex Pillow imageio imageio-ffmpeg
pip install git+https://github.com/waitingcheung/diffusers.git@feat/motif-video
import torch
from diffusers import (
AdaptiveProjectedGuidance,
DPMSolverMultistepScheduler,
MotifVideoPipeline,
)
from diffusers.utils import export_to_video
guider = AdaptiveProjectedGuidance(
guidance_scale=8.0,
adaptive_projected_guidance_rescale=12.0,
adaptive_projected_guidance_momentum=0.1,
use_original_formulation=True,
normalization_dims="spatial",
)
pipe = MotifVideoPipeline.from_pretrained(
"Motif-Technologies/Motif-Video-2B",
torch_dtype=torch.bfloat16,
guider=guider,
)
pipe = pipe.to("cuda")
output = pipe(
prompt="A woman standing in a sunlit field as flower petals swirl around her in slow motion. Each petal floats gently through the golden light, casting tiny shadows. Her hair moves like water, and time seems to stand still.",
negative_prompt="text overlay, graphic overlay, watermark, logo, subtitles, timestamp, broadcast graphics, UI elements, random letters, frozen pose, rigid, static expression, jerky motion, mechanical motion, discontinuous motion, flat framing, depthless, dull lighting, monotone, crushed shadows, blown-out highlights, shifting background, fading background, poor continuity, identity drift, deformation, flickering, ghosting, smearing, duplication, mutated proportions, inconsistent clothing, flat colors, desaturated, tonally compressed, poor background separation, exposure shift, uneven brightness, color balance shift",
height=736,
width=1280,
num_frames=121,
num_inference_steps=50,
frame_rate=24,
)
export_to_video(output.frames[0], "output.mp4", fps=24)
import torch
from diffusers import (
AdaptiveProjectedGuidance,
DPMSolverMultistepScheduler,
MotifVideoPipeline,
)
from diffusers.utils import export_to_video, load_image
guider = AdaptiveProjectedGuidance(
guidance_scale=8.0,
adaptive_projected_guidance_rescale=12.0,
adaptive_projected_guidance_momentum=0.1,
use_original_formulation=True,
normalization_dims="spatial",
)
pipe = MotifVideoImage2VideoPipeline.from_pretrained(
"Motif-Technologies/Motif-Video-2B",
torch_dtype=torch.bfloat16,
guider=guider,
)
pipe = pipe.to("cuda")
image = load_image("https://huggingface.co/Motif-Technologies/Motif-Video-2B/resolve/main/assets/i2v_sample.jpg")
output = pipe(
prompt="Three friends stride through a sun-bleached meadow as a warm breeze ripples the tall dry grass around their legs.",
negative_prompt="text overlay, graphic overlay, watermark, logo, subtitles, timestamp, broadcast graphics, UI elements, random letters, frozen pose, rigid, static expression, jerky motion, mechanical motion, discontinuous motion, flat framing, depthless, dull lighting, monotone, crushed shadows, blown-out highlights, shifting background, fading background, poor continuity, identity drift, deformation, flickering, ghosting, smearing, duplication, mutated proportions, inconsistent clothing, flat colors, desaturated, tonally compressed, poor background separation, exposure shift, uneven brightness, color balance shift",
image=image,
height=736,
width=1280,
num_frames=121,
num_inference_steps=50,
frame_rate=24,
)
export_to_video(output.frames[0], "output.mp4", fps=24)
# Text-to-Video (default settings)
python inference.py \
--prompt "A woman standing in a sunlit field as..." \
--output t2v_output.mp4
# With SageAttention (~2x faster, requires sageattention package)
python inference.py \
--prompt "Three friends stride through a sun-bleached meadow..." \
--use-sage-attention \
--output t2v_output.mp4
See inference.py --help for all available options.
| Parameter | Default | Notes |
|---|---|---|
| Resolution | 1280×736 | 720p, best quality |
| Frames | 121 | ~5 seconds at 24fps |
| Scheduler | DPMSolver++ | solver_order=2, flow_shift=15.0 |
| Guidance scale | 8.0 | With APG (normalization_dims="spatial") |
| Inference steps | 50 | |
| Negative prompt | (built-in) | See code examples above |
use_linear_quadratic_schedule | False | Must be set explicitly |
| dtype | bfloat16 | Recommended for H100/A100 |
For GPUs with 24 GB or less (e.g. RTX 4090, RTX 3090), CPU offloading and FP8 quantization can reduce peak VRAM from ~30 GB to ~15 GB with minimal speed impact.
| Mode | Peak VRAM | Recommended GPU |
|---|---|---|
pipe.to("cuda") | ~30 GB | A100, H100, H200 |
enable_model_cpu_offload() | ~19 GB | RTX 4090, RTX 3090 |
+ FP8 quantization | ~15 GB | RTX 4090, RTX 3090 |
Full guide → docs/memory-efficient-inference.md
GGUF quantized weights at Motif-Video-2B-GGUF — up to 2.7 GB VRAM savings with no speed penalty. Combined with SageAttention for ~1.6× faster inference.
| Variant | Sage (s/it) | Speedup | Peak alloc (GB) |
|---|---|---|---|
| BF16 | 14.75 | 1.58x | 15.12 |
| Q8_0 | 14.49 | 1.60x | 13.44 |
| Q4_K_M | 14.59 | 1.60x | 12.53 |
Full guide → docs/gguf-sageattention.md
Official ComfyUI custom nodes: ComfyUI-MotifVideo2B
Note: Currently requires High VRAM mode. GGUF quantized model loading in ComfyUI is in progress.
Motif-Video 2B achieves the highest Total Score among open-source models we evaluate.
| Model | Params | Total | Quality | Semantic |
|---|---|---|---|---|
| Wan2.2-T2V (prompt-opt.) | A14B | 84.23 | 85.42 | 79.50 |
| Motif-Video 2B (Ours) | 2B | 83.76 | 84.59 | 80.44 |
| SANA-Video | 2B | 83.71 | 84.35 | 81.35 |
| Wan2.1-T2V | 14B | 83.69 | 85.59 | 76.11 |
| OpenSora 2.0 (T2I2V) | 11B | 83.60 | 84.40 | 80.30 |
| Wan2.1-T2V | 1.3B | 83.31 | 85.23 | 75.65 |
| HunyuanVideo | 13B | 83.24 | 85.09 | 75.82 |
| CogVideoX1.5-5B (prompt-opt.) | 5B | 82.17 | 82.78 | 79.76 |
| Step-Video-T2V | 30B | 81.83 | 84.46 | 71.28 |
| LTX-Video | 2B | 80.00 | 82.30 | 70.79 |
Notable per-dimension highlights for Motif-Video 2B (open-source):
The full 16-dimension breakdown is in Table 3 of the technical report.
A note on VBench vs. perceptual quality. Motif-Video 2B leads on VBench Total Score, but in our internal side-by-side comparisons against Wan2.1-T2V-14B we observe a perceptual gap in favor of the larger model on temporal stability and fine human anatomy. We discuss the sources of this gap (uniform dimension weighting, near-correct semantic credit) in Section 7 of the report. We report the gap explicitly rather than smoothing it over.
In a blind pairwise study against six contemporaneous open-source baselines (SANA-Video, LTX-Video 2, Wan2.1-14B, Wan2.1-1.3B, Wan2.2-5B, CogVideoX-5B) on 40 LLM-generated prompts, Motif-Video 2B is preferred over both SANA-Video (similar parameter count) and Wan2.1-1.3B (similar parameter count, larger training corpus) on prompt-following and video-fidelity axes. Wan2.1-14B remains the preferred model overall, consistent with its 7× larger parameter count and substantially larger training data.
We report limitations as the boundary conditions under which the design decisions in this report should be interpreted, not as caveats.
We view temporal stability and data coverage — not architectural depth — as the primary remaining ceilings on this model. Both are the most natural axes for a future iteration that the current architecture is built to absorb.
If you find Motif-Video 2B useful in your research, please cite:
@techreport{motifvideo2b2026,
title = {Motif-Video 2B: Technical Report},
author = {Motif Technologies},
year = {2026},
institution = {Motif Technologies},
url = {https://arxiv.org/abs/2604.16503}
}
We build on a number of excellent open-source projects, including the Wan2.1 VAE [Wan Team, 2025], T5Gemma / Gemma 3 [Google], TREAD [Krause et al., 2025], REPA with the V-JEPA family of visual encoders [Bardes et al.], DDT [Wang et al.], and the broader diffusers and Accelerate ecosystems. Compute was provisioned on Microsoft Azure and orchestrated with SkyPilot on Kubernetes.
This model is released under the Apache 2.0 License. See LICENSE for details.