Downloads · 30 days
0
AkaneTendo25/Cosmos3-ConvRot
Cosmos3-ConvRot is a image-to-video model from AkaneTendo25. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Weight-only quantized transformers of the NVIDIA Cosmos3 models, for use with ComfyUI-Cosmos3. Each file is one quantized transformer; take the VAE, tokenizer(s) and config.json from the matching official nvidia/Cosmo…
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2026
Repo size
373 GB
Likes
2
Public
Click a slice to open those files.
.safetensors373 GB · 100%
From the Hugging Face model README
Weight-only quantized transformers of the NVIDIA Cosmos3 models, for use with
ComfyUI-Cosmos3. Each file is one quantized
transformer; take the VAE, tokenizer(s) and config.json from the matching official
nvidia/Cosmos3-* repo.
Two formats:
Both are weight-only: activations stay bf16, so this lowers memory/download size, not compute speed. int4 runs at about bf16 speed — the per-forward dequantize and un-rotate add a little.
| Model | bf16 | int8 | int4 |
|---|---|---|---|
Cosmos3-Nano | 30 GB | 16.5 GB | 12.4 GB |
Cosmos3-Super | 128 GB | 65.7 GB | 46.8 GB |
Cosmos3-Super-Image2Video | 128 GB | 65.6 GB | 46.7 GB |
Cosmos3-Super-Image2Video-4Step | 128 GB | 65.6 GB | 46.7 GB |
Cosmos3-Edge | 6.7 GB | 3.9 GB | 3.0 GB |
File names: Cosmos3-<name>-int8-convrot.safetensors and Cosmos3-<name>-int4-convrot.safetensors.
int4 and int8 are provided for every model.
Prompting note (Edge): Cosmos3-Edge is trained on JSON-structured prompts and is less robust to
plain text than the larger Nano/Super. Plain text usually works, but on some detailed scenes (notably
reflective surfaces) it can produce flare/pulsation artifacts; wrapping the text as
{"temporal_caption": "<your prompt>"} avoids them. This is a base-model property, not a quantization
effect (it shows in bf16 too).
With ComfyUI's dynamic VRAM the transformer is streamed from host RAM, so the GPU holds only the activations. Measured at 832×480, 93 frames, at the minimum VRAM budget (maximum streaming):
| Model | Min VRAM | RAM (bf16 / int8 / int4) |
|---|---|---|
Cosmos3-Edge | ≈6 GB | 14 / 7 / 7 GB |
Cosmos3-Nano | ≈7 GB | 58 / 21 / 20 GB |
Cosmos3-Super (t2v & i2v) | ≈8–9 GB | 240 / 67 / 63 GB |
Min VRAM is the activation floor (set by resolution × frame count, not the weight format). RAM is the peak host memory — larger than the file on disk (staging + overhead), and bf16 peaks near twice the weight size. RAM and VRAM trade off: giving the GPU more VRAM holds more weights on-card and lowers the RAM figure. The Super-family int4 checkpoints fit a 64 GB host (≈63 GB); int8 needs a little more (≈67 GB).
nvidia/Cosmos3-<name> into ComfyUI/models/cosmos3/<name>/.transformer/ folder, delete the bf16 shards and *.index.json, then put the quantized
file there renamed to diffusion_pytorch_model.safetensors. Keep the official config.json,
vae/, text_tokenizer/, sound_tokenizer/.weight_dtype = default); the loader reads the format from the
checkpoint metadata. Requires comfy-kitchen (int8 from ComfyUI >= 0.27) and the latest
ComfyUI-Cosmos3 (int4 needs the ConvRot-aware loader).Reproducible in method, not bit-for-bit (the calibration set and RNG vary per run). Both formats keep
the same escape set in bf16: proj_in, proj_out, time_embedder, audio_proj,
modality_embed, the embeddings, and all norms and biases.
Weight-only symmetric INT8, per-output-channel scale (weight_scale, float32, [out, 1]), with a
group-wise Hadamard rotation (ConvRot, group 256) applied before quantization and undone at load.
comfy_quant tag int8_tensorwise, convrot=true. No calibration: at 8-bit the per-channel scale and
the rotation keep the round-to-nearest error small — error feedback (GPTQ) is only needed at int4.
Produced with
convert_to_quant:
ctq -i transformer_bf16.safetensors -o out_int8_convrot.safetensors \
--int8 --scaling_mode row --simple --convrot --convrot-group-size 256 \
--comfy_quant --save-quant-metadata --cosmos3 --device cuda --low-memory
Round-to-nearest INT4 — even with the ConvRot rotation — leaves visible artifacts on these models, so the MLP path uses GPTQ error compensation on real activations. Steps:
uni_pc_bh2; 4-step model: 4 steps, cfg 1, euler) over
≈4 prompts, with a forward pre-hook on every target linear. Keep a reservoir of up to 4096 rows
per layer (random replacement beyond that). The understanding tower sees the text prefill; the
generation tower sees every denoising step.H = XᵀX · 2/N from the rotated activations; diagonal damping raised through
{0.01, 0.03, 0.1, 0.3, 1, 3} × mean(diag) until the Cholesky factors; columns processed in blocks
of 128 with per-column error feedback into the not-yet-quantized columns; plain round-to-nearest
only if damping never succeeds.[out, 1] scale). INT4 on attention produces visible artifacts.Layer counts follow the checkpoint's config.json — e.g. Cosmos3-Super-Image2Video packs 384 MLP
linears (INT4) + 512 attention linears (INT8).
Execution. Both formats run as an explicit dequantize-then-bf16 matmul (unpack the 4-bit weights
to bf16, un-rotate, multiply in bf16). This is weight-only regardless: there is no int4-weight ×
bf16-activation tensor-core op on any GPU (Hopper or Blackwell — Blackwell's FP4 cores are for NVFP4,
both operands 4-bit, not this W4A16 layout), so a W4A16 kernel would dequantize internally too, with no
compute speedup either way. We dequantize explicitly instead of calling comfy-kitchen's
gemv_awq_w4a16, which is non-deterministic (atomic accumulation jitters the video frame-to-frame) and
numerically off on these shapes.
Derived from NVIDIA Cosmos3 checkpoints; the OpenMDW-1.1 license
applies (same as the upstream nvidia/Cosmos3-* models).