Downloads · 30 days
2
1% of all-time downloads
wfen/Cosmos3-Nano-FP8-Blockwise
Cosmos3-Nano-FP8-Blockwise is a machine learning model from wfen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as openmdw-1.0.
Mixed-precision blockwise FP8 weight-only quantization of Cosmos3OmniTransformer.
Downloads · 30 days
2
1% of all-time downloads
All-time downloads
134
Public
Parameters
15.2B
42.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors23.9 GB · 100%
How the weights are stored.
F8_E4M310.9B · 72%
From the Hugging Face model README
Mixed-precision blockwise FP8 weight-only quantization of Cosmos3OmniTransformer.
Checkout Cosmos3-Nano-WebUI to setup a quick web interface for inferencing images/videos/texts/actions with this checkpoint. Screenshots and walkthrough are available in this document.
# Verify the checkpoint loads correctly
python load_checkpoint.py --verify
# Generate a single frame (smoke test)
python load_checkpoint.py \
--prompt "A robotic arm in a kitchen" \
--steps 8 --frames 1
# Generate a multi-frame video (quality)
python load_checkpoint.py \
--prompt "A robotic arm in a kitchen" \
--steps 35 --frames 57 \
--height 480 --width 640
All available in the project's Docker environment.
transformer/
config.json # Cosmos3OmniTransformer config (action_gen=False)
diffusion_pytorch_model.safetensors # FP8 weights + blockwise scales (~18.8 GB)
modelopt_state.pt # Structural sidecar (~670 KB, quantizer topology)
quantization_config.json # Recipe, block size, scale layout documentation
quantizer_map_diff.json # INV-2 validation result
load_checkpoint.py # Standalone loader
README.md # This file
The modelopt_state.pt sidecar contains only the quantizer structure (which modules have quantizers, their configs). It does NOT contain model weights. It uses pickle format (weights_only=False on load) and should only be trusted from this locally-produced checkpoint.
import torch, glob
import modelopt.torch.opt as mto
from diffusers import Cosmos3OmniPipeline, Cosmos3OmniTransformer, UniPCMultistepScheduler
from safetensors.torch import load_file
CKPT = "dist/Cosmos3-Nano-FP8-Blockwise"
# 1. Build skeleton
cfg = {**Cosmos3OmniTransformer.load_config(f"{CKPT}/transformer/config.json"), "action_gen": False}
transformer = Cosmos3OmniTransformer.from_config(cfg).to(torch.bfloat16)
# 2. Restore quantizer structure from sidecar
state = torch.load(f"{CKPT}/transformer/modelopt_state.pt", weights_only=False)
restored = mto.restore_from_modelopt_state(transformer, state)
if restored:
transformer = restored
# 3. Load weights + scales from safetensors
tensors = {}
for shard in sorted(glob.glob(f"{CKPT}/transformer/*.safetensors")):
tensors.update(load_file(shard))
transformer.load_state_dict(tensors, strict=True)
# 4. Build pipeline
pipe = Cosmos3OmniPipeline.from_pretrained(
CKPT, transformer=transformer, torch_dtype=torch.bfloat16, enable_safety_checker=False
)
pipe.scheduler = UniPCMultistepScheduler.from_config(pipe.scheduler.config, flow_shift=10.0)
pipe = pipe.to("cuda")
# 5. Generate under autocast
with torch.autocast("cuda", torch.bfloat16):
result = pipe(prompt="...", num_frames=57, height=480, width=640, num_inference_steps=35,
generator=torch.Generator("cpu").manual_seed(123))
Compared to Phase 1 per-tensor FP8 (vs bf16 gold standard):
| Case | Improved? | LPIPS Delta |
|---|---|---|
| EC-01 (t2v) | No | -0.056 |
| EC-02 (sound/MoE) | Yes | +0.069 |
| EC-03 (i2v) | Yes | +0.010 |
| EC-05 (hard) | Yes | +0.029 |
| EC-06 (OOD) | Yes | +0.025 |
4/6 cases improved (DoD-3 PASS). See docs/reports/phase_2_quality_report.md for the full comparison.
For reproducible output, use these exact settings:
UniPCMultistepScheduler(flow_shift=10.0)CUBLAS_WORKSPACE_CONFIG=:4096:8torch.autocast("cuda", torch.bfloat16)"cpu" (not "cuda")In folder 'assets/FP8-Examples', you can find a selection of generated videos from the FP8 blockwise checkpoint. Each example includes VRAM reports and ffprobe reports.