Downloads · 30 days
0
SyFeee/LTX2.3-Dual-Character-en
LTX2.3-Dual-Character-en is a image-to-video model from SyFeee. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A field-tested image-to-video character-consistency LoRA for Lightricks/LTX-2.3 (22B distilled), tuned for two-character dialogue scenes and multi-shot cinematic video generation.
Downloads · 30 days
0
Access
Public
Updated May 21, 2026
Repo size
332 MB
Likes
20
Public
Click a slice to open those files.
.safetensors327 MB · 99%
From the Hugging Face model README
A field-tested image-to-video character-consistency LoRA for Lightricks/LTX-2.3 (22B distilled), tuned for two-character dialogue scenes and multi-shot cinematic video generation.
⚠️ Naming note (corrected 2026-05-21): The original filename and ModelScope repo include the string "IC-LORA", but this is NOT an IC-LoRA in the strict technical sense (parallel-canvas /
video_conditioningmechanism). An A/B/C test (same prompt + seed, three reference-channel variants) confirmed that the LoRA's actual conditioning mechanism is first-frame pixel pinning (the regular i2v path), not parallel-canvas attention. Earlier copy on this card incorrectly described it as IC-LoRA — that has been removed. Credit to ZKong for raising the discrepancy in the discussions tab.
Episode is an 8-shot Chinese palace drama (《玉佩定情》 + 《暗夜阴谋》) with three characters: 沈月华 (Shen Yuehua, heroine), 萧云霄 (Xiao Yunxiao, prince), 慕容静 (Murong Jing, antagonist). Render config: 1280×704, 121 frames @ 24 fps, ambient audio.
<video controls autoplay muted loop src="https://huggingface.co/SyFeee/LTX2.3-Dual-Character-en/resolve/main/examples/E1S1_garden_walk_single_character.mp4"></video>
<video controls autoplay muted loop src="https://huggingface.co/SyFeee/LTX2.3-Dual-Character-en/resolve/main/examples/E1S2_prince_meets_dual_character.mp4"></video>
<video controls autoplay muted loop src="https://huggingface.co/SyFeee/LTX2.3-Dual-Character-en/resolve/main/examples/E2S1_murong_plots_cross_scene.mp4"></video>
<video controls autoplay muted loop src="https://huggingface.co/SyFeee/LTX2.3-Dual-Character-en/resolve/main/examples/E2S4_three_character_confrontation.mp4"></video>
Fine-tuned on Lightricks/LTX-2.3 (22B distilled), specifically for:
The reference image is consumed via first-frame pixel pin (standard i2v conditioning), not via the parallel-canvas / video_conditioning channel.
# Upstream LTX-2.3 distilled pipeline — single reference as first-frame pin
from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.args import ImageConditioningInput
from ltx_core.loader import LoraPathStrengthAndSDOps, sd_ops as _sd_ops_mod
import torch
lora = LoraPathStrengthAndSDOps(
"LTX2.3-IC-LORA-Dual-Character.safetensors",
0.8, # strength (standalone)
_sd_ops_mod.LTXV_LORA_COMFY_RENAMING_MAP,
)
pipe = DistilledPipeline(
distilled_checkpoint_path="ltx-2.3-22b-distilled-1.1.safetensors",
spatial_upsampler_path="ltx-2.3-spatial-upscaler-x2-1.1.safetensors",
gemma_root="google/gemma-3-12b-it-qat-q4_0-unquantized",
loras=[lora],
device=torch.device("cuda:0"),
)
video, audio = pipe(
prompt="...",
seed=42,
height=704, width=1280,
num_frames=121, # 5 s @ 24 fps, satisfies 8k+1
frame_rate=24,
images=[ImageConditioningInput( # first-frame pin = THE reference mechanism
path="character_ref.png",
frame_idx=0,
strength=0.9,
)],
enhance_prompt=False,
)
LTX's i2v pin rejects two pins at the same frame_idx, so two refs can't both be pinned at frame 0. Two workable patterns:
Pattern A (recommended): composite reference image. Build one image with character A on the left and character B on the right (e.g., via PIL Image.paste or any image editor), pin THAT at frame_idx=0. Both identities transfer in one pin.
Pattern B: stagger the pins. Pin character A at frame 0, character B at a later latent boundary (e.g., frame 64 — must be a multiple of 8 per the VAE's temporal compression). Only works if B doesn't need to be visible from the very first frame.
| Setting | Value |
|---|---|
| Resolution | 1280 × 704 (16:9, native LTX-2.3 distilled training resolution) |
| Faster preview | 960 × 544 (~40% faster, slightly less detail) |
| Frames | satisfy 8k+1 — e.g. 121 (5 s), 193 (8 s), 241 (10 s), 361 (15 s) at 24 fps |
| Strength | Standalone 0.7-0.9 · stacked with style LoRAs 0.3-0.5 |
| Pin strength | 0.85-0.95 for tight identity, 0.7 for looser "inspired-by" |
| Trigger word | None |
Quirks of this LoRA + the LTX-2.3 distilled backbone that aren't in the original card but matter in practice.
This LoRA has a light-wuxia-robe bias. Dark outfits drift toward white at low pin strength. Repeat the color token glued to each clothing noun:
BAD: black fedora and black suit
GOOD: BLACK fedora, white shirt, BLACK suit jacket, BLACK trousers,
... BLACK suit, BLACK trousers throughout
Also bump pin strength to ~0.95 for color fidelity on dark outfits.
This LoRA was trained on Chinese drama clips with burned-in Chinese subtitles. Any quoted dialogue (「…」 or "…") in the prompt causes the LoRA to hallucinate subtitle characters at the bottom of the frame. Single biggest gotcha.
BAD: 低声警告 「此茶不可饮!」 ← fake on-screen subtitles
GOOD: 低声急切警告她茶水有毒 ← clean output, indirect narration
If your app needs subtitles, burn them post-hoc via ffmpeg drawtext.
At high motion intensity, the model loses object tracking. A directive like "fedora flies off mid-spin and tumbles to the floor" produces broken output — the hat dematerialises. Either:
For multi-shot dialogue scenes, character identity drifts across cuts. Workaround: re-pin the reference image at frame 0 of every shot. (Deterministic seed + same first-frame pin + same prompt scaffolding produces good repeatability.)
On consumer hardware (RTX 4090 24 GB), expect ~3-4 minutes per shot.
The original Chinese model card from ModelScope is reproduced below for users who want the unmodified original documentation. (Note: the original card uses the "IC-LoRA" label — the term has been kept here for fidelity, even though the A/B/C test described above shows the conditioning mechanism is first-frame i2v pinning rather than parallel-canvas IC-LoRA.)
<details> <summary>点击展开原版中文模型卡片 (click to expand original Chinese README)</summary>本模型是基于 Lightricks LTX-2.3 底模训练的 IC-LoRA,专为双人同框对话、角色互动及分镜头视频生成场景深度优化。
一、模型核心提升
二、模型基本信息
三、运行指南
四、推荐参数配置
五、Prompt 编写规范
六、效果说明与局限性
| GPU | VRAM | Works? |
|---|---|---|
| A100 / A800 80 GB | 80 GB | ✅ ~70 s per 5 s shot |
| RTX 4090 / 3090 | 24 GB | ✅ ~3-4 min per 5 s shot |
| RTX 4080 / 4070 Ti Super | 16 GB | ❌ won't fit 22B in bf16 |
| anything < 24 GB | — | ❌ no |
This is an English-language mirror of fxj1131's LTX2.3 Dual-Character LoRA on ModelScope. All credit for the model weights belongs to the original author, 麻雀 AI (Maque AI). This mirror exists to make the model + documentation accessible to HuggingFace users who cannot easily access ModelScope, and to share field-tested usage notes from a production deployment. The
.safetensorsweights file is unmodified and byte-identical to the ModelScope upload.
Apache License 2.0 — same as the original. See LICENSE and NOTICE.