Downloads · 30 days
1.1K
84% of all-time downloads
Smite79/MiniMax-H3-Longvideos
MiniMax-H3-Longvideos is a text-to-video model from Smite79. Use it when you need video from a text prompt. It is set up for minimax-h3. The card lists the license as other.
Downloads · 30 days
1.1K
84% of all-time downloads
All-time downloads
1.3K
Public
Repo size
28.1 MB
Likes
136
Public
Click a slice to open those files.
.safetensors28.1 MB · 98%
From the Hugging Face model README
This node is a constant work in progress! If you are noticing bugs or features that do not work, please ensure that you are pulling the most recent version and updating your workflows.
Please note that RealRebelAI has been blacklisted from this project. If you want further updates for the node, please continue to use my updates as the node is being continously worked on. Don't support those who steal other people's work for their own credit.
Long MiniMax-H3 video with synchronised audio from a single prompt, in ComfyUI.
H3 renders up to about 15 seconds at a time. This node renders your scene shot by shot, starts each shot on the last frame of the one before, and joins them into one video with one soundtrack. Your text reaches the model word for word.
Copy this folder into ComfyUI/custom_nodes/ and restart the ComfyUI server.
Requires ComfyUI with native MiniMax-H3 support.
Updating from an earlier version: the node's widgets have changed. Right-click the
node → Fix node (recreate), set your values again and save the workflow. Inputs an
old workflow still has linked that the node no longer uses are ignored and named in
info.
| UNET | a MiniMax-H3 model — a hybrid fl2va/ref2va merge works best |
| CLIP | H3's text encoder, loader type minimax |
| VAE | the H3 video VAE |
| audio VAE | the H3 audio VAE — a separate file, and it must be the converted one |
UNETLoader ─┐ images ─> Video Combine / Save Video
CLIPLoader ─┼─> H3-LongVideos ─> audio ─┘
VAELoader ──┘ info ─> Show Text
prompt is an input socket — wire a multiline text node into it.
Turn on plan_only to see the exact text every shot will get, without rendering.
Paragraphs are separated by a blank line.
anchor (optional) is framing for the whole film: look, camera, lighting,
location. It goes at the front of every shot. When it is filled in, every paragraph
of the prompt is a shot and there is no scene paragraph.character_memory (optional) is who is in the film and what they wear. It
follows the scene in every shot.A dim cell with a cot. Mara <Picture 1> is a tall woman in a grey dress. Dan <Picture 2> is a guard in a dark uniform.
Mara sits on the cot and stares at the door.
Dan handcuffs her wrists behind her back.
hold: Mara, handcuffs behind her back
Dan presses duct tape over her mouth.
hold: Mara, duct tape over her mouth
Dan says "Not a sound." He walks out.
The node reads restraints and gags from your beats and carries them from shot to shot: handcuffs, zip ties, rope, chains, shackles, duct tape, gags and ball gags, blindfolds and collars. It reads them being put on (Dan handcuffs her wrists behind her back, Dan grabs Mara's wrists and cuffs them, Dan takes out the handcuffs and snaps them onto her wrists, Dan wraps duct tape around her mouth), already on (her wrists cuffed behind her back, Mara sits in handcuffs, Mara sits with duct tape over her mouth), and coming off (Dan pulls the tape off, Dan unlocks the cuffs).
info instead of guessing.
Add a hold: line for those.info
as not read. If they belong on someone, add a hold: line.info instead.info says what goes on, what is held, and what comes off.The node reads clothes coming off and going back on (Max takes off her jacket,
Crystal pulls her thong down her legs, Crystal puts her jacket back on). From the
next shot on, a garment that came off is taken out of that person's description in
the scene paragraph, the anchor and character_memory, so no shot asks for it again.
It comes back once they put it on.
info.These lines are taken out of the text before the model sees it. hold: and release:
correct the restraint reading, and remove: and wear: the clothing reading. For the
person they name, they replace whatever was read from that beat.
| line | what it does |
|---|---|
hold: Name, item; item | Name is wearing or bound with these items. From the next shot on, every prompt states Name: item; item., that it all stays on for the whole shot, and where held wrists or ankles stay, until the item comes off. |
release: Name, item | the item comes off in this shot. release: Name takes everything off Name. |
seconds: 6 | this shot's length, whatever shot_length says. |
remove: Name, garment | the garment comes off in this shot and leaves Name's description from the next shot on. |
wear: Name, garment | the garment goes back on, and back into the description from the next shot on. |
exit: Name | Name leaves during this shot and is gone from the next. A named exit such as Dan walks out is read without it; use it for anything else, such as He leaves. |
cut | this shot starts fresh instead of on the previous shot's last frame. Use it for a new place or time. |
A hold: line in the scene paragraph, the anchor or the character memory means the
person starts the video already held.
The node also reads these lines when they are written on the same line as the
sentence, as their own paragraph, with a dash or colon after the name, or without the
comma. A line it still cannot read, such as a hold with no name, is listed in info.
In a hold: line, write each item the way it should look: handcuffs behind her back,
wrists zip tied in front, rope around her ankles, duct tape over her mouth.
The shot where an item goes on is described by your own sentence, and it closes with
the item in plain view at its last frame. The item joins the held line from the shot
after, so it is not drawn before the action happens.
The node keeps track of who is present, so people appear only when they should.
character_memory (Mara: a tall woman), a
name written right before a picture tag (Mara <Picture 1>), or a name on a hold:
or exit: line.cut, only the people that beat names are present.For someone who is not present:
character_memory line is left outWhen everyone in a shot is a declared character, a shot with one or two people also
says how many are in it. The count is left out whenever someone else might be there:
a name that is not declared, or a guard, the man, a crowd. Declare every
character, with a character_memory line or a picture tag, so the guarding covers
them all.
Every shot that starts on the previous shot's last frame says so in its prompt: the same place and the same people, one moment earlier, with nobody new joining. That frame reaches H3 as a picture, and a picture the prompt does not mention is read as another person.
After a cut, someone who was last seen alone is given that frame as their current
look, in place of their portrait. Their clothes and anything held on them carry over.
H3 comes out a little more saturated and contrasty each time it continues from a frame,
and left alone that burns the picture more with every shot. The first shot, and the
first shot after a cut, are left as rendered and set the look. Every shot that
continues from a frame is graded where it picks up: the first frame you see is matched
to the frame it continues from (brightness, contrast, saturation and the shape of the
shadows and highlights), since both show the same moment, and that same correction is
carried through the rest of the shot. Nothing later in a shot is forced onto another
frame's tones, so someone stepping in, the lights going down or the camera turning keep
their own look. A continued shot that opens washed out gets its colour back the same way.
A black frame, or a shot that does not pick up where the last one ended, is left as
rendered. Each continued shot's numbers (how its opening came out against the frame it
continues from, whether it was graded, and its finished look against shot 1) go to the
ComfyUI log.
Wire up to four pictures into ref_image_1 … ref_image_4 and refer to them in the text
as <Picture 1> … <Picture 4>. A tag with no picture wired is removed.
Once a person is held, their own picture (written right after their name, as in
Mara <Picture 1>) is left out of shots that start on the previous frame. A portrait
without the cuffs or the tape pulls them back off. After a cut the picture comes back,
unless a later frame of them stands in for it.
The pictures work on every checkpoint and LoRA. The hybrid b25-49 (fl2va_ref2va) and any
checkpoint whose file name says ref2va were trained on them and get them in every shot.
FL2VA checkpoints, such as the base minimax_h3_fl2va and 10Eros, and FastH3 were not:
when a person is already in the frame a shot opens on and their picture comes in as well,
they draw that person twice. So on those, a person's picture is left out of a shot that
continues from a frame they are in, and that frame carries their look. The first shot,
the shots after a cut, and anyone walking into a continued shot still get their
picture. The node reads which checkpoint is loaded from its own model link.
"…", '…', “…”, ‘…’ or <d>…</d>)
speaks. Every line reaches H3 inside <d>…</d>, the speech tags it was trained on,
with the words exactly as written. Its first half second is held quiet so the line
does not start on the cut.shot_seconds, up to H3's 15 seconds, so words are not rushed or cut off.
A seconds: line still wins. A line too long for one shot is flagged in info: split
it across beats.foley_resolution to full for a full-size pass. A beat that describes
its own sounds keeps those.silence_wordless off to
render every shot in one pass with free audio instead.ambient_audio to have it looped under the
whole soundtrack at ambient_level. It plays under the model's sound and does not
change it.Text alone cannot stop H3, or a LoRA, from freeing restrained arms or catching a fall with cuffed hands. Every shot where someone's arms or ankles are held is checked with DWPose after it renders, in one of two ways.
With the H3 Fun controlnet: the skeleton is held, on any H3 checkpoint and LoRA. The
node loads minimax_h3_fun_controlnet_union_pruned_int8_convrot.safetensors from
models/model_patches by itself, or uses the one wired into pose_controlnet.
pose_end of the steps. If the draft
gives nothing to hold, the shot is rendered at full size without it.models/diffusion_models; only its small curve is read.Without the controlnet: the shot is rendered again. If the restraint broke, the
shot is rendered again on a new seed, up to pose_retries more times. The node keeps
the take where it held, or else the take where it broke in the fewest frames. Set
pose_retries to 0 to turn this off.
Needs: the DWPose files from comfyui_controlnet_aux:
ckpts/hr16/yolox-onnx/yolox_l.torchscript.ptckpts/hr16/DWPose-TorchScript-BatchSize5/dw-ll_ucoco_384_bs5.torchscript.ptWhat is checked:
hold: lines:
behind her back, in front, above her head, at her waist, ankles, ankles to
her wrists.to the bed)info says which way the restraints are being kept and has a line for every checked
shot.latent_upscale samples each shot at resolution and enlarges it in latent space
before decoding, which is much cheaper than sampling large.
models/latent_upscale_models.latent_upscale_scale sets the factor.upscale enlarges the finished video:
rtx: NVIDIA RTX Video Super Resolution (needs the comfyui_nvidia_rtx_nodes pack)model: the model picked in upscale_model, from models/upscale_modelslanczos: a plain resizeupscale_target_short_edge fits the result's short edge to that many pixels.
lanczos needs it; for rtx it also sets the factor.cudaMallocAsync
allocator, restart ComfyUI with --disable-cuda-malloc.
euler). Leave sigmas unwired.
hyperflow_endpoint_v1.0.safetensors
beside the node, or from the original minimax_h3_hyperflow_8step_v1.0.safetensors
in models/loras.info says so.| widget | |
|---|---|
resolution, megapixels | aspect preset and size; at 1.0 each preset is its native size, 0 keeps it exactly |
shot_seconds | the longest a shot may be, and every shot's length when shot_length is fixed |
shot_length | from the beat sizes each shot from its own beat: about 2.2s per action, or the spoken line, plus one action's time where something is put on, capped by shot_seconds. fixed gives every shot shot_seconds |
steps, sampler_name, scheduler, seed | as in KSampler; one seed for the whole video |
first_frame | the first shot starts on this picture |
sigmas | your own schedule (overrides Hyperflow's grid) |
shift_video, shift_audio | H3's sigma shift, 12/3 by default |
silence_wordless | keep shots without a quoted line silent |
plan_only | report the shots and their text without rendering |
pose_strength, pose_end | how strongly, and for how much of the schedule, the skeleton holds |
pose_retries | without pose control, how many times a shot whose restraint broke is rendered again |
anchor | framing at the front of every shot; filled in, every paragraph is a shot |
character_memory | who is in the film, after the scene in every shot |
latent_upscale, latent_upscale_scale | upscale each shot in latent space, and by how much |
upscale, upscale_model, upscale_target_short_edge | upscale the finished video |
ambient_audio, ambient_level | an audio bed looped under the whole soundtrack, and how loud |
foley_resolution | the size of the picture copy the sound pass listens to: half (about six times faster) or full |
| slot | what it is |
|---|---|
images | the finished frames |
audio | the synchronised soundtrack |
info | what the node did for each shot, how long it took, and its contrast and colour against shot 1 |
script | the exact text each shot was given |
frames_per_shot, total_frames, shots, video_seconds | for downstream nodes |
models/diffusion_models; info says how many.megapixels or
shot_seconds.The owner of this repo will not be responsible for any copyright strikes incurred because of use. You are responsible for your works. Use this node responsibly and ethically.