Downloads · 30 days
0
SpXMerlin1D/MiniMaxH3-CondBridge-Qwen3.5-4B
MiniMaxH3-CondBridge-Qwen3.5-4B is a text-to-video model from SpXMerlin1D. Use it when you need video from a text prompt. It is set up for adapter.
A 1.14B adapter that bridges a lightweight Qwen3.5-4B student to the MiniMax-H3 33B text encoder's injection space. It converts the student's hidden states into the exact CLIP-injection representation the H3 DiT expec…
Downloads · 30 days
0
Access
Public
Updated Aug 16, 2026
Repo size
2.3 GB
Likes
7
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
A 1.14B adapter that bridges a lightweight Qwen3.5-4B student to the
MiniMax-H3 33B text encoder's injection space. It converts the student's
hidden states into the exact CLIP-injection representation the H3 DiT expects
(post-condition_proj + token_refiner), replacing the 33B teacher encoder
end-to-end.
MiniMax-H3 video generation conditions the DiT on text embeddings produced by a
33B text encoder (condition_proj + 2-layer token_refiner). Running it locally
is heavy. CondBridge distills that interface into a 1.14B adapter that
consumes:
| Input | Shape | Source |
|---|---|---|
h3_ids | [S_T] | H3 tokenizer (vocab 151,643) |
student_hidden | [S_S, 2560] | Qwen3.5-4B hidden_states[-1] (post-final-norm) |
and outputs the teacher-equivalent representation [1, S_T, 5376]. Because the
output lives in the same space as the teacher's post-refiner embeddings, the DiT
consumes it directly
H3Adapter — 1.144B params total:
| Module | Params | Role |
|---|---|---|
source_projection | 13.8M | student 2560 → 5376 (KV) |
query_embedding | 40.3M | H3 ids → 5376 query (embed 151,936×256 + proj) |
cross_attention | 319M | 32-head resampler, QK-norm + tanh gate |
token_refiner | 751M | 2 pre-norm blocks + final RMSNorm (mirrors teacher) |
Forward: kv = source_proj(student) → x = cross_attn(query_embed(h3_ids), kv)
→ token_refiner(x).
condition_proj + token_refiner output of the official encoder)Two-stage fine-tuning (32GB GPU, bf16):
| Stage | Duration | Steps | Scope | Result |
|---|---|---|---|---|
| 1 | 4h | 18,874 | adapter body, refiner frozen (DiT init) | cos 0.8856 |
| 2 | 1.5h | 4,135 | all params, refiner lr×0.1 | cos 0.9229 |
huber×1.0 + cos×0.8 + infonce×0.05 + sp×0.1 + mag×0.1 + stat×0.0002
with a curriculum ramp on the contrastive termspn = normalize(pred) × target_norm)Held-out 1,002 prompts (disjoint from train):
| Metric | Value |
|---|---|
| cosine (token-level) | 0.9229 |
| MSE | 0.8162 |
| norm_ratio (pred/target scale) | 1.013 |
import torch
from transformers import AutoTokenizer
from adapter.model import H3Adapter # repo code
from safetensors.torch import load_file
adapter = H3Adapter().to(torch.bfloat16)
adapter.load_state_dict(load_file("condbridge.safetensors"))
adapter.eval()
h3_tok = AutoTokenizer.from_pretrained("<h3 tokenizer>")
# student_hidden: Qwen3.5-4B hidden_states[-1] [S_S, 2560]
h3_ids = h3_tok(prompt, add_special_tokens=False)["input_ids"]
embeds = adapter(torch.tensor([h3_ids]), student_hidden.unsqueeze(0)) # [1, S_T, 5376]
Requires: Qwen3.5-4B (student), MiniMax-H3 tokenizer, the H3 DiT.
Use with the [ComfyUI-MiniMaxH3-Adapter] node: MiniMaxH3AdapterLoader
(student dir + adapter .safetensors) → plug into official
MiniMaxH3ImageToVideo. The node exposes the adapter as a duck-typed CLIP.
** I use HauhauCS/Qwen3.5-4B-Uncensored-HauhauCS-Aggressive but it worked for Qwen/Qwen3.5-4B and every Fine-Tune or Quantizations **
一个 1.14B 参数的适配器,为轻量级 Qwen3.5-4B 学生模型搭建通往 MiniMax-H3
33B 文本编码器注入空间的"桥"。 它将学生模型的隐藏状态转换为 H3 DiT 期望的
CLIP 注入表示(condition_proj + token_refiner 之后的结果),端到端替代
33B 教师编码器。
MiniMax-H3 视频生成用 33B 文本编码器(condition_proj + 2 层 token_refiner)
产生文本嵌入来条件化 DiT。本地运行它很重。CondBridge 蒸馏了这套接口,用
一个 1.14B 适配器消费:
| 输入 | 形状 | 来源 |
|---|---|---|
h3_ids | [S_T] | H3 tokenizer(词表 151,643) |
student_hidden | [S_S, 2560] | Qwen3.5-4B hidden_states[-1](final-norm 后) |
并输出与教师等价的表示 [1, S_T, 5376]。由于输出与教师的 post-refiner 嵌入
处于同一空间,DiT 可直接消费
H3Adapter — 共 1.144B 参数:
| 模块 | 参数量 | 作用 |
|---|---|---|
source_projection | 13.8M | 学生 2560 → 5376(作为 KV) |
query_embedding | 40.3M | H3 ids → 5376 查询(embed 151,936×256 + proj) |
cross_attention | 319M | 32 头重采样器,QK-norm + tanh 门控 |
token_refiner | 751M | 2 个 pre-norm block + final RMSNorm(镜像教师) |
前向:kv = source_proj(student) → x = cross_attn(query_embed(h3_ids), kv)
→ token_refiner(x)。
condition_proj + token_refiner 输出)两阶段微调(32GB 显卡,bf16):
| 阶段 | 时长 | 步数 | 范围 | 结果 |
|---|---|---|---|---|
| 1 | 4h | 18,874 | 适配器主体,refiner 冻结(DiT 初始化) | cos 0.8856 |
| 2 | 1.5h | 4,135 | 全参数,refiner lr×0.1 | cos 0.9229 |
huber×1.0 + cos×0.8 + infonce×0.05 + sp×0.1 + mag×0.1 + stat×0.0002,
对比项带课程 ramppn = normalize(pred) × target_norm)Held-out 1,002 条 prompt(与训练集不重叠):
| 指标 | 数值 |
|---|---|
| cosine(逐 token) | 0.9229 |
| MSE | 0.8162 |
| norm_ratio(预测/目标尺度) | 1.013 |
import torch
from transformers import AutoTokenizer
from adapter.model import H3Adapter # 仓库代码
from safetensors.torch import load_file
adapter = H3Adapter().to(torch.bfloat16)
adapter.load_state_dict(load_file("condbridge.safetensors"))
adapter.eval()
h3_tok = AutoTokenizer.from_pretrained("<h3 tokenizer>")
# student_hidden: Qwen3.5-4B hidden_states[-1] [S_S, 2560]
h3_ids = h3_tok(prompt, add_special_tokens=False)["input_ids"]
embeds = adapter(torch.tensor([h3_ids]), student_hidden.unsqueeze(0)) # [1, S_T, 5376]
依赖:Qwen3.5-4B(学生)、MiniMax-H3 tokenizer、H3 DiT。
配合 [ComfyUI-MiniMaxH3-Adapter] 节点使用:MiniMaxH3AdapterLoader
(学生目录 + 适配器 .safetensors)→ 接入官方 MiniMaxH3ImageToVideo 节点。
** 我在测试时使用HauhauCS/Qwen3.5-4B-Uncensored-HauhauCS-Aggressive,但理论上原版Qwen/Qwen3.5-4B的GGUF量化以及其任何微调的GGUF量化版本都可用 **