Downloads · 30 days
83
92% of all-time downloads
rockstarengine/vflash
vflash is a image-to-video model from rockstarengine. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for minimax-h3. The card lists the license as other.
This is a small experimental adapter for MiniMax H3 Base FL2VA, not a base model, not a low-step distillation adapter, and not the retired Ref2VA INT8 checkpoint. Use the published weights at scale −1.0, with the offi…
Downloads · 30 days
83
92% of all-time downloads
All-time downloads
90
Public
Repo size
34.1 GB
Likes
0
Public
Click a slice to open those files.
.safetensors65.6 MB · 100%
From the Hugging Face model README
This is a small experimental adapter for MiniMax H3 Base FL2VA, not a base model, not a low-step distillation adapter, and not the retired Ref2VA INT8 checkpoint. Use the published weights at scale −1.0, with the official 16-step schedule. Positive scale +1.0 is the original training direction and degraded quality in our pilot.
We observed reduced local mesh artifacts, double contours and some improved structural clarity in a limited set of animation and live-action examples. Fast-motion artifacts remain, and face identity, physical interaction, prompt adherence and audio have not passed a general quality qualification. This is not a universal bad-frame fix.
| Item | Value |
|---|---|
| Base | MiniMaxAI/MiniMax-H3, FL2VA/Base transformer (not Ref2VA) |
| Base revision | 42ed227ee7df40d41602854ae760620d6eb651fe |
| File | h3-base16-reverse-lora-v1.safetensors |
| Size | 65,627,696 bytes |
| Layout | DiffSynth H3 / PEFT; fused per-head interleaved QKV |
| Contents | 208 FP32 tensors; 50 DiT blocks and 2 TokenRefiner blocks |
| Modules | attn.qkv_proj, attn.out_proj |
| Rank / alpha | 8 / 8 |
| Parameters | 16,400,384 |
| Inference multiplier | −1.0, applied to B only or via an equivalent runtime scale |
| Training | 12 clips, 64.125 seconds, 96 single-clip updates, learning rate 5e−5 |
| Components | PyTorch / PEFT / DiffSynth 7686e54d41d25c0e8ed5f1318acc23b6bb832654 |
The file is the unchanged saved training checkpoint, not a newly trained or negated file. Each clip was seen eight times with changing noise. Two work-disjoint validation clips and additional held-out generation controls were used; the full available caption corpus was not used. The small mixed animation/CG/live-action pilot is not representative of all video tasks. Training media, captions, user references and private prompts are not distributed here; this release grants no rights to any underlying media.
The original checkpoint metadata includes training_only=true and
quality_qualified=false. These accurately preserve its experimental origin and are
not a claim that it is a production-qualified adapter. No conversion or hashing of the
base model is required to use this release.
Download adapter.json, this card, LICENSE, NOTICE and the adapter; pin a Hub commit
revision for reproducibility:
hf download rockstarengine/vflash h3-base16-reverse-lora-v1.safetensors adapter.json README.md LICENSE NOTICE --revision <commit> --local-dir h3-reverse-lora-v1
The following shows the exact adapter loading contract, assuming dit is an already
loaded, unadapted H3 transformer with the pinned DiffSynth module layout. It is not a
complete video pipeline. Keep the official tokenizer, conditioning, VAE and schedulers.
import torch
from peft import LoraConfig, inject_adapter_in_model
from safetensors.torch import load_file
state = load_file("h3-reverse-lora-v1/h3-base16-reverse-lora-v1.safetensors")
assert len(state) == 208
assert all(k.startswith("pipe.dit.") and v.dtype == torch.float32 for k, v in state.items())
assert not any("lora_" in k for k, _ in dit.named_parameters())
dit = inject_adapter_in_model(
LoraConfig(r=8, lora_alpha=8, target_modules=["attn.qkv_proj", "attn.out_proj"]),
dit,
)
for name, parameter in dit.named_parameters():
if "lora_" in name:
parameter.data = parameter.data.float()
signed = {
k.removeprefix("pipe.dit."): (-v if ".lora_B." in k else v)
for k, v in state.items()
}
result = dit.load_state_dict(signed, strict=False)
assert not result.unexpected_keys
assert not any("lora_" in k for k in result.missing_keys)
dit.requires_grad_(False)
This snippet folds −1 into B, so do not also set a runtime scale of −1. Do not negate both A and B: that cancels the sign change. Keep LoRA parameters and residual arithmetic in FP32; casting the whole adapted model to BF16 or merging into BF16 weights changes the tested numerical contract. Standard Diffusers/ComfyUI loaders cannot be assumed to understand this fused per-head QKV format. A layout conversion must split Q/K/V B rows per head and duplicate A, while preserving both TokenRefiner blocks.
The vflash 0.5.6 source adds explicit Python/CLI loading to the complete pipeline:
from pathlib import Path
from vflash.pipeline import AttentionAdapter, H3Pipeline
with H3Pipeline(
prepared, device=device, trust_local_code=True,
attention_adapter=AttentionAdapter(
Path("h3-reverse-lora-v1/h3-base16-reverse-lora-v1.safetensors"),
rank=8, scale=-1.0,
),
) as pipeline:
result = pipeline.generate(request, Path("result.mp4"))
The prepared assets and device must select single-SM89 official Base16. The same owner
handles conditioning, denoising and MP4 delivery, loads the adapter once, and can serve
serial requests. Each result records the DiT-only scope, rank, FP32 residual precision
and signed scale. Omit attention_adapter to restore unadapted behavior on a new owner.
For vflash generate, supply --attention-adapter, --attention-adapter-rank 8 and
--attention-adapter-scale -1 together. No automatic download or HTTP option is added.
Full-pipeline GPU qualification is tracked separately from the existing native evidence.
Vflash v0.5.4 adds an explicit experimental low-level context. With an already active, exclusively owned, single-SM89 Base16 native session:
from pathlib import Path
from safetensors.torch import load_file
from vflash.native.h3_attention_lora import apply_dit_attention_lora
state = load_file("h3-reverse-lora-v1/h3-base16-reverse-lora-v1.safetensors", device="cpu")
with apply_dit_attention_lora(session.runtime, state, rank=8, scale=-1.0):
result = session.generate(Path("bundles/example"), Path("outputs/reverse-latents.safetensors"))
This uses the unmodified published checkpoint, so -1 is applied once in the context, not by also negating B. It selects the 50 DiT blocks and intentionally does not update TokenRefiner. It is therefore distinct from the full DiffSynth example above. Use block-ring for Sol, serial requests and official 16 NFE. Output is AV latents; use the matching official decoder for video delivery. Exit restores the original methods without merging or changing base weights. See the complete interface and limits. There is no automatic loading or change to serving defaults. The complete pipeline option above is newer than this low-level context; HTTP configuration remains unchanged.
The weights are a MiniMax H3 model derivative governed by the accompanying MiniMax H3
Community License Agreement, not the Apache-2.0 license of the vflash source package.
Read the full license, including territorial, redistribution, acceptable-use and
commercial conditions, before accessing or using this adapter. Its Applicable Territory
excludes the EU, UK, Republic of Korea and USA. Downloading or using the derivative is
subject to those terms; public Hub visibility does not waive them. See NOTICE.
This release replaces the obsolete Ref2VA INT8 files and their instructions in the repository's current tree. Historical commits remain for traceability; they are not recommended downloads. No training media, private examples or service configuration are included. Powered by MiniMax H3.