Downloads · 30 days
39
30% of all-time downloads
007ILOVEU/LTX2.3-ICEdit-Insight
LTX2.3-ICEdit-Insight is a video-to-video model from 007ILOVEU. Use it for the video-to-video task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as apache-2.0.
<table <tr <td align="center"<img src="./assets/effects/output004.webp" alt="Video restoration preview" width="420"/</td <td align="center"<img src="./assets/effects/视频高清对比效果.webp" alt="Video HD enhancement preview" w…
Downloads · 30 days
39
30% of all-time downloads
All-time downloads
128
Public
Repo size
30.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors30.2 GB · 100%
From the Hugging Face model README
LTX2.3-ICEdit-Insight is a task-aware video restoration and editing model family developed by JoyFox Lab, built on top of the LTX-2.3 DiT-based audio-video foundation model.
This release focuses on four practical video editing directions:
Unlike conventional frame-level enhancement pipelines, this model family operates as a generative video restoration system in latent video space. It is designed to preserve global structure, camera motion, object identity, and temporal consistency while reconstructing missing or degraded visual content.
Project links: GitHub project | JoyFox on Hugging Face | Paper (Research Square) | DOI
10.21203/rs.3.rs-9775063/v1This model release corresponds to the paper's unified video post-processing framework built around three high-level settings: video super-resolution/enhancement, occlusion removal and repair, and instruction-driven semantic editing.
In the currently released inference package, those ideas are exposed through four practical routes:
The paper introduces three core components:
| File | Purpose |
|---|---|
ltx-2.3-edit-insight-dev-fp8.safetensors | Unified Insight base checkpoint for LTX-2.3 editing |
ltx2.3-video-restoration-general.safetensors | Video restoration, artifact cleanup, blur and noise recovery |
ltx2.3-ic-video-upscale-general.safetensors | Video HD enhancement, super-resolution, and detail recovery |
ltx2.3-ic-watermark-remove-general.safetensors | Watermark removal and occlusion-aware reconstruction |
ltx2.3-ic-subtitles-remove-general.safetensors | Subtitle removal and text overlay cleanup |
Run all scripts from the project root.
bash run_restoration.sh
bash run_hd.sh
bash run_hd.sh /path/to/input.mp4
bash run_watermark_rm.sh
bash run_watermark_rm.sh /path/to/input.mp4
bash run_subtitle_rm.sh
bash run_subtitle_rm.sh /path/to/input.mp4
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
python run_pipeline.py \
--mode restoration \
--video ./inputs/input_480p.mp4 \
--prompt "Convert the video to ultra-high-definition quality while removing artifacts and rebuilding high-frequency details." \
--output ./outputs/output_restoration.mp4 \
--height 1184 --width 704 --num-frames 97 \
--fps 24.0 --seed 42 \
--sigma-profile workflow \
--streaming-prefetch-count 2 \
--model-checkpoint ./models/checkpoints/ltx-2.3-edit-insight-dev-fp8.safetensors \
--lora ./models/loras/ltx2.3-train/ltx2.3-video-restoration-general.safetensors
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
python run_pipeline.py \
--mode hd \
--video ./inputs/input_480p.mp4 \
--prompt "Convert the video to ultra-high-definition quality, significantly improving clarity, fine detail richness, texture fidelity, and overall perceptual sharpness." \
--output ./outputs/output_hd.mp4 \
--height 1184 --width 704 --num-frames 97 \
--fps 24.0 --seed 42 \
--sigma-profile workflow \
--streaming-prefetch-count 2 \
--model-checkpoint ./models/checkpoints/ltx-2.3-edit-insight-dev-fp8.safetensors \
--lora ./models/loras/ltx2.3-train/ltx2.3-ic-video-upscale-general.safetensors
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
python run_pipeline.py \
--mode watermark_rm \
--video ./inputs/input_480p.mp4 \
--prompt "Remove short-video platform watermarks and related occlusions from the video, restoring a clean, clear, and natural original image." \
--output ./outputs/output_watermark_rm.mp4 \
--height 1184 --width 704 --num-frames 97 \
--fps 24.0 --seed 1546 \
--sigma-profile workflow \
--streaming-prefetch-count 2 \
--model-checkpoint ./models/checkpoints/ltx-2.3-edit-insight-dev-fp8.safetensors \
--lora ./models/loras/ltx2.3-train/ltx2.3-ic-watermark-remove-general.safetensors
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
python run_pipeline.py \
--mode subtitle_rm \
--video ./inputs/input_480p.mp4 \
--prompt "Remove subtitles, captions, and related text occlusions from the video, restoring a clean and natural underlying image." \
--output ./outputs/output_subtitle_rm.mp4 \
--height 1184 --width 704 --num-frames 97 \
--fps 24.0 --seed 42 \
--sigma-profile workflow \
--streaming-prefetch-count 2 \
--model-checkpoint ./models/checkpoints/ltx-2.3-edit-insight-dev-fp8.safetensors \
--lora ./models/loras/ltx2.3-train/ltx2.3-ic-subtitles-remove-general.safetensors
We introduce a task-aware IC-Edit training framework for LTX-2.3, where each restoration direction is optimized with dedicated instruction conditioning and task-specific IC-LoRA adapters.
The model is trained not only to improve visual quality, but also to understand the editing goal behind different restoration tasks, including watermark removal, subtitle cleanup, damaged region recovery, and high-definition enhancement.
The model family is built on the LTX-2.3 foundation architecture, a diffusion-transformer video model designed for high-fidelity image-to-video and video generation workflows.
Our adaptation targets video restoration by improving:
Video restoration requires more than strong single-frame quality. We optimize temporal consistency so that restored areas remain stable across adjacent frames.
This reduces common artifacts such as:
The training curriculum covers realistic video defects including:
This improves generalization across short videos, social-media clips, mobile footage, downloaded videos, and compressed production material.
For watermark and subtitle removal, the model is optimized to reconstruct the hidden visual content behind occluded regions.
Instead of smearing or blurring the target area, it uses surrounding spatial context and temporal cues to infer plausible background structure, object boundaries, lighting, and texture continuity.
For HD enhancement, the model improves perceptual sharpness and fine visual detail through frequency-aware restoration training.
This is especially helpful for recovering:
8k + 1 rule.32 in single-stage inference.If this model family is useful in your research or application workflow, please cite:
@article{tang2026ltxinsight,
title = {LTX-Insight: Unified Video Restoration and Semantic Editing via Task-Aware Adaptation and Temporal Consistency},
author = {Tang, Fan and Li, Siyuan},
journal = {Research Square},
year = {2026},
month = {05},
doi = {10.21203/rs.3.rs-9775063/v1},
url = {https://www.researchsquare.com/article/rs-9775063/v1}
}
This model family was trained and optimized by JoyFox Lab (Chengdu Xuanhu Technology Co., Ltd.).
The training pipeline includes:
For research collaboration, commercial licensing, or workflow integration, contact:
Licensed under Apache 2.0.
Please also review the license terms of the upstream LTX-2.3 base model when using or redistributing derivative checkpoints.