Downloads · 30 days
0
topabaem/LTX-2.5-Text-Encoder-4bit-8GB
LTX-2.5-Text-Encoder-4bit-8GB is a image-to-video model from topabaem. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Downloads · 30 days
0
Access
Public
Updated Aug 22, 2026
Repo size
17.1 GB
Likes
1
Public
Click a slice to open those files.
.safetensors16.9 GB · 99%
From the Hugging Face model README
The 8 GB in the name is the file, not the card. On disk this is 8.46 GB against the original's 26.264 GB. Running it needs about 9.7 GB of VRAM — an 8 GB card is not enough. See Memory.
EN — I'm a student researching ML quantization. Getting this one done burned through so much in server bills that from here on I'll only be able to afford niu lai movies. Thank you for using the model.
한국어 — 저는 ML 양자화를 연구하는 학생입니다. 양자화를 진행하면서 서버 비용을 너무 많이 써서, 앞으로 영화는 niu lai만 봐야 할 것 같습니다. 모델을 사용해 주셔서 감사합니다.
中文 — 我是一名研究机器学习量化的学生。做这次量化烧掉了太多服务器费用, 以后看电影大概只能看 niu lai 了。感谢您使用这个模型。
☕ Buy me a coffee · 커피 한 잔 사주기 · 请我喝杯咖啡
The Gemma4-12B text encoder that LTX-2.5 needs in order to read a prompt, compressed from 26.264 GB to 8.46 GB (3.10x) and runnable on any CUDA GPU — no minimum compute capability, no custom kernels, no CUDA 13.
If you have been unable to run LTX-2.5 because the text encoder alone wanted 26 GB, this is the part that was in your way. Drop it in and the rest of the model is unchanged.
Built from the encoder Lightricks published on 2026-08-17
(1b92891c, "Aligns the published encoders with the LTX-2.5 model
checkpoints") — the revision that matches the released DiT.
The packed model files have not changed. The public Space is deployed at
257e2908c4d60e8a190244b9d7c6a4d625444ec7 with conditioning contract
gemma4-raw-intermediate-slots-v1: raw intermediate hidden states, the
model-returned final normalized slot, BOS token 2, left padding and masking to
1024 tokens, physical position IDs, and a valid-token tail crop.
Four authenticated two-second smoke renders on that exact repository and runtime SHA produced valid 1024x640, 24 fps, 49-frame clips with audio. The tested prompts contained 40, 61, 258, and 260 tokens including BOS. All four passed the file/stream checks, but none passed every requested audiovisual criterion. The telephone clips kept the requested subject, composition, and background, yet the handset deformed or overlapped during motion and visible movement did not reliably align with repeated ringing. The short forge clip kept an anvil, hammer, and hot metal but omitted the blacksmith and recognizable horseshoe and produced four strikes rather than three. The long forge prompt changed to an unrelated outdoor person. This proves that the corrected path runs; it does not prove that motion stays artifact-free or that long prompts are reliably followed. No corrected-contract BF16 render oracle has been published yet.
Every other public quantization of this encoder needs recent hardware:
| build | size | needs |
|---|---|---|
| Lightricks BF16 | 26.264 GB | — |
Lightricks comfy-int8-convrot | 15.373 GB | cc 8.9 |
DmitryDB nvfp4 | 11.197 GB | cc 8.9, comfy_kitchen, CUDA 13 |
joeygambino / Winnougan w4a8 | 10.604 GB | SM 8.0+ |
| this | 8.46 GB | nothing beyond PyTorch |
Dequantization happens on the CPU at load and the resident model is BF16, so there is no kernel requirement to satisfy. Verified running on Tesla V100 (cc 7.0), A100 (cc 8.0) and RTX 6000 Ada (cc 8.9).
It is also the smallest of the set, because it quantizes the three tensors the
other nvfp4 recipe protects as BF16 "precision islands" — embed_tokens and
both aggregate tables, 4.4 GB of the source.
| file | build | video relL2 | audio relL2 | ||Q||/||W|| |
|---|---|---|---|---|
A3.packed.safetensors | bypass guard + group-bounded AWQ | 0.05204 | 0.04827 | 4.216 |
A0.packed.safetensors | legacy guard, for comparison | 0.06061 | 0.06095 | 302.654 |
Use A3. A0 is published only so the comparison can be checked; it
carries weights up to 300x their proper norm in near-dead channels, which is
harmless on this calibration set and fragile by construction.
samples/ holds every clip twice — once from the BF16 original (26.264 GB)
and once from this 4-bit build, with everything downstream of the encoder
held identical: same DiT, same seed, same schedule, same VAE settings, one
process. compare-NN.mp4 stacks each pair, BF16 on the left.
Historical preprocessing notice. Every sample below predates the 2026-08-22 conditioning correction. Those renders applied the encoder's learned final norm to all 49 hidden-state slots; the current path preserves raw intermediate slots and keeps only the final model-returned slot normalized. Some old sample paths also differed in padding. The paired BF16/4-bit clips are still controlled comparisons within that legacy preprocessing, but they are not a current-runtime oracle. The packed weights themselves did not change.
Two sets, because they answer different questions. compare-NN.mp4 uses the
vendor's euler_ancestral, which is what you actually get; it re-rolls noise
every step, so the two builds return different takes. compare-det-NN.mp4
uses deterministic euler, where the seed fixes the starting noise and the
conditioning is the only thing left that can move a pixel — that is the set
that attributes a difference to the encoder.
Mean absolute error against BF16 over every decoded frame:
| clip | euler_ancestral | euler |
|---|---|---|
| robot in rain | 0.0172 | 0.0182 |
| dune | 0.0596 | 0.0580 |
| forge | 0.0098 | 0.0129 |
| night road | 0.0702 | 0.0252 |
| smoke | 0.0128 | 0.0165 |
Four of the five deterministic pairs hold together — same composition, same lighting, same timing, differing in surface detail. The dune does not: BF16 renders a soldier in fatigues where the 4-bit renders a man in a business suit, from the same seed under a deterministic sampler. That difference belongs to the encoder, and it is published rather than cropped out. The qualification it deserves is that neither build followed that prompt — it asked for an astronaut and got neither — so the model had no confident answer there for a small conditioning change to disturb.
All five deterministic pairs at a glance — BF16 left, 4-bit right, one row per prompt. Rows 1, 3, 4 and 5 hold together; row 2, the dune, is where the compression is visible.

<video controls width="100%" poster="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/moe-idol-poster.png" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/moe-idol-15s.mp4"></video>

361 frames with sound, from a 300-word prompt naming ten sequenced gestures, made
entirely through an earlier revision of the Space
on ZeroGPU in 198 s. It is a historical sample, not proof of the current
conditioning contract or general long-prompt compliance. The full prompt is in
samples/idol/prompt.txt.
Within that earlier runtime, two other pipeline defaults had to be fixed before a clip this long held together, and both were in this project's own defaults rather than in the compression:
DISTILLED_SIGMAS is nine
fixed numbers; LTXVScheduler derives its shift from the latent's token count,
and a long clip carries several times more. Left fixed, the sample never
converges — furniture renders semi-transparent and saturation halves. Clips
past 15 s now switch automatically.Details, the measurements, and the failures behind both are in
samples/long/schedule.md.

Six clips at 1024x640 — sample at 512x320, upscale the latent 2x, sample
again. That second pass had been skipped here from the start on an assumption
that a 16 GB card could not afford it; measured, it peaks at 10.03 GiB, and it
is worth 4.1x the Laplacian variance. Everything else on this page is
one-pass and softer for it. samples/sharp/, including
the one of the six that does not follow its prompt and why the earlier padding
attribution is now considered confounded.
<video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/idol/idol-15s.mp4"></video>
353 frames at 512x320, generated in 123 s at a 6.26 GiB peak — no more memory
than the ten-second version, so the length ceiling was not reached. The prompt
and an honest account of what it did not do — the audio is not speech, and the
"2D anime" instruction is ignored at other lengths by the BF16 original too —
are in samples/idol/.
Forge — the pair holding together. Deterministic sampler, same seed. Left is BF16, right is 4-bit.
<video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/compare-det-02.mp4"></video>
Dune — the pair that does not. Same terrain, same sun, same shadow, same walk, and a different person.
<video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/compare-det-01.mp4"></video>
Individually, the forge under the vendor's own sampler:
| BF16 original | 4-bit |
|---|---|
| <video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/bf16-02.mp4"></video> | <video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/4bit-02.mp4"></video> |
samples/README.md carries the prompts, the per-clip conditioning drift, and
what these clips do and do not establish. All thirty clips are in samples/,
and the Space plays them side by side under its BF16 vs 4-bit tab.
| resident (default) | dequantized | |
|---|---|---|
| PyTorch allocated, peak | 8.48 GiB | ~22.3 GiB |
| PyTorch reserved, peak | 9.33 GiB | — |
what nvidia-smi shows | 9.70 GiB | — |
Size your card from the last row. The first is max_memory_allocated, which
counts only live allocator blocks — it misses the CUDA context and everything
the caching allocator reserved and has not handed back, and it undercounts by
nearly 2 GiB here. Earlier versions of this card quoted that number, and anyone
who bought an 8 GB card on the strength of it would have been wrong.
Measured on a Tesla V100-SXM2-16GB (cc 7.0) with a 55-token prompt, through the ComfyUI-faithful path that left-pads to 1024 tokens. So:
The two modes produce torch.equal conditioning, so the choice is footprint
against speed and never quality. Resident dequantizes inside forward, which
costs time on short prompts; dequantized builds one dense BF16 model at load and
then needs a card that can hold 26 GB.
These figures are for the text encoder alone. Generating video also needs the
DiT and the VAEs, which this repository does not contain — the pipeline in the
Space loads a Q3_K_M DiT alongside it.
Five packages and one file. No build step, no custom CUDA kernels, no compilation.
pip install -r <(curl -sL https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/requirements.txt)
from huggingface_hub import hf_hub_download, snapshot_download
repo = "topabaem/LTX-2.5-Text-Encoder-4bit-8GB"
# The loader ships with the weights; put it on the path before importing it.
import sys, os
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "ltx_packed_codec.py")))
from ltx_packed_codec import load_packed_model
from transformers import AutoTokenizer
packed = hf_hub_download(repo, "A3.packed.safetensors")
encoder_dir = snapshot_download(repo, allow_patterns=["encoder-hf/*"]) + "/encoder-hf"
model = load_packed_model(encoder_dir, packed, resident=True)
tokenizer = AutoTokenizer.from_pretrained(encoder_dir)
Verified end to end in a clean virtualenv containing nothing but those five
packages, on torch 2.13.0 / transformers 5.15.1 and on torch 2.10.0 /
transformers 5.12.1. encoder-hf/ is config and tokenizer only, 31 MB — the
26 GB original is not needed.
Nothing unusual. The format needs no fp8 hardware: the group scales are stored
as float8_e4m3fn bytes and converted in software during a CPU-side decode, so
float8 here is a container and never an instruction. There is no minimum
compute capability, no comfy_kitchen, no CUDA 13. The resident model is BF16
and runs on cards with no bf16 tensor cores at all.
The one real constraint is your torch wheel, not your GPU. The default wheel on PyPI is now a cu130 build and cu130 dropped Volta. Measured on a V100:
| torch build | device | result |
|---|---|---|
| 2.13.0+cu130 | CPU | works |
| 2.13.0+cu130 | V100, sm_70 | no kernel image is available |
| 2.10.0+cu128 | V100, sm_70 | works |
That failure arrives at the first kernel launch, well after
torch.cuda.is_available() has returned True, so it does not look like an
installation problem — which is why the loader checks first. Before reading
a byte of the 8.46 GB it compares your card against
torch.cuda.get_arch_list() and, on a mismatch, stops with the fix:
this torch (2.13.0+cu130) has no kernels for Tesla V100-SXM2-16GB (sm_70).
It was built for sm_75, sm_80, sm_86, sm_90, sm_100, sm_120, and the first CUDA
op would fail with 'no kernel image is available for execution on the device'.
The model is fine - it needs no custom kernels. Install a torch built for your
card, e.g. for sm_70:
pip install torch --index-url https://download.pytorch.org/whl/cu128
or pass device='cpu' to load without touching the GPU.
On sm_70 install a cu128 build:
pip install torch --index-url https://download.pytorch.org/whl/cu128
Ampere and newer are unaffected — the stock wheel carries kernels for them.
This repository currently ships the standalone packed loader
ltx_packed_codec.py; it does not include a ComfyUI custom-node package.
The public Space uses a separate ComfyUI-backed generation runtime and the packed
encoder, but that does not make this model repository a drop-in ComfyUI node.
Space: LTX-2.5 Text Encoder 4bit
— text-, image- and video-to-video, running this encoder on ZeroGPU. Measured
there at 512x320, 25 frames: t2v 53.0 s, i2v 46.9 s, v2v 37.6 s. The current
runtime SHA is 257e2908c4d60e8a190244b9d7c6a4d625444ec7; see
Current Space status for the bounded
post-deployment result rather than assuming every long prompt will be followed.
Both work, and both needed a fix ComfyUI does not ship. LTXVAddGuide is the
only producer of guided LTX latents and it cannot take an LTX-2.5 one: it calls
torch.cat on what is a NestedTensor for this model. The ValueError in that
function saying AV guides are unsupported never fires — NestedTensor.shape
proxies to the video half, whose channel count is exactly the 128 it checks for
— so the real failure is a TypeError, and the message is stale.
Everything below the node already supports AV guides: the model routes
keyframe_idxs to the video branch, the sampler pads a video-only denoise mask
with ones for audio, and model_base splits the packed mask apart again. So
ltx_av_guide.py unwraps the pair, runs the stock node on the video half and
re-wraps; the guide arithmetic stays the vendor's.
Measured on a 16 GB V100 at 512x320, 25 frames: t2v 98.8 s / 5.84 GiB, i2v 62.7 s / 6.80 GiB, v2v 86.5 s / 5.84 GiB. An i2v first frame lands relL2 0.0766 from its guide image against 0.7061 for the same seed and prompt without the guide, so the guide is honoured rather than merely accepted.
nvfp4 4.5 bpw (E2M1, group 16 with an fp8-e4m3 scale) on 320 projections;
int8 row-wise on embed_tokens and both aggregate tables; norms and asset blobs
left BF16. Per-tensor AWQ alpha search, then sequential GPTQ error compensation
(blocksize 128, percdamp 0.01), packed inside the build because the group scales
cannot be recovered afterwards.
A3 adds three things the plain recipe lacks:
mean(diag(H)) — the relative
rule moves with the calibration Hessian's dynamic range, which made the old
guard's best setting differ between Volta and Ada.x·diag(1/s) @ Q(W·diag(s))ᵀ holds for any s, so this needs no format
change.A3 needed a 10x
escalation, written into the artifact metadata so a damped build is never
silently compared with an undamped one.verify: all 686
tensors value-exact. No drift here is attributable to the format.A3 is
closer on four of five (mean 0.05931 against A0's 0.06571).gemma4-raw-intermediate-slots-v1 contract is also
not yet available.Source Lightricks/LTX-2.5 revision 1b92891c+, torch 2.11.0+cu128,
transformers 5.14.1, plan r45c, calibration calib-large.txt. Evidence in
evidence/: gate JSON, build logs with per-layer drift and ||Q||/||W||, the
A100 BF16 reference, all three conditionings, and fifteen rendered clips.
Earlier builds against the pre-2026-08-17 encoder, and the method results that
came from them, are at topabaem/Pacific-LTX-2.5-Encoder-r45d.