Downloads · 30 days
13
3% of all-time downloads
openaiarka/JoyAI-VL-Interaction-Preview-AWQ
JoyAI-VL-Interaction-Preview-AWQ is a machine learning model from openaiarka. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
AWQ W4A16 post-training quantization (PTQ) of jdopensource/JoyAI-VL-Interaction-Preview.
Downloads · 30 days
13
3% of all-time downloads
All-time downloads
438
Public
Parameters
8.8B
7.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors7.2 GB · 100%
How the weights are stored.
I326.9B · 79%
From the Hugging Face model README
AWQ W4A16 post-training quantization (PTQ) of jdopensource/JoyAI-VL-Interaction-Preview.
This model card focuses on how the quantized checkpoint was produced, not on how to consume the model (for inference usage please refer to the base model card).
| Attribute | Value |
|---|---|
| Base model | jdopensource/JoyAI-VL-Interaction-Preview |
| Architecture | Qwen3VLForConditionalGeneration |
| Scale | ~8B parameters |
| Original dtype | bfloat16 |
| Attribute | Value |
|---|---|
| Method | AWQ — Activation-aware Weight Quantization |
| Tool | llmcompressor |
| API used | llmcompressor.oneshot() with AWQModifier |
| Quantization scheme | W4A16 (4-bit weights, 16-bit activations) |
| Format | compressed-tensors (quant_method: "compressed-tensors") |
| Quantization status | compressed |
| Group size | 128 |
| Weight packing | pack-quantized (INT4 packed into INT32 containers) |
| Duo scaling | Enabled (duo_scaling=True) |
| AWQ grid search size | n_grid: 20 |
| Symmetric weights | Yes |
| Smoothing mappings | Auto-inferred by llmcompressor for Qwen3VLForConditionalGeneration |
Linear layers in the text model.lm_headvisual, vision_tower, vision_model, vision_proj, mergerThe vision encoder and language model head were intentionally kept in higher precision to preserve visual understanding quality and output embedding fidelity.
The exact excluded module list recorded in config.json is:
["model.visual.blocks.0.attn.qkv", "model.visual.blocks.0.attn.proj", ...]
(fully expanded in the repository's config.json under quantization_config.ignore)
recipe.yaml)default_stage:
default_modifiers:
AWQModifier:
mappings:
- smooth_layer: re:.*input_layernorm$
balance_layers: ['re:.*q_proj$', 're:.*k_proj$', 're:.*v_proj$']
- smooth_layer: re:.*v_proj$
balance_layers: ['re:.*o_proj$']
- smooth_layer: re:.*post_attention_layernorm$
balance_layers: ['re:.*gate_proj$', 're:.*up_proj$']
- smooth_layer: re:.*up_proj$
balance_layers: ['re:.*down_proj$']
duo_scaling: true
n_grid: 20
QuantizationModifier:
targets: [Linear]
ignore: [lm_head, 're:.*visual.*', 're:.*vision_tower.*', 're:.*vision_model.*',
're:.*vision_proj.*', 're:.*merger.*']
scheme: W4A16
bypass_divisibility_checks: false
llmcompressor internally split the request into an AWQ smoothing step followed by a QuantizationModifier step.
Because the original jdopensource/JoyAI-VL-Interaction dataset only contains annotation JSON files and not the actual video assets, calibration was performed on a publicly available video-question-answering dataset that ships with real .mp4 files.
| Attribute | Value |
|---|---|
| Dataset | MBZUAI/VCGBench-Diverse |
| Videos available | 877 unique .mp4 files |
| QA pairs available | 4,354 entries in vcgbench_diverse_qa.json |
| Samples used | 128 QA pairs |
| Sampling seed | 42 |
| Filtering rule | Only QA entries whose referenced video file actually exists were kept |
For each sampled QA pair:
.mp4 with OpenCV.RGB image.[
{"type": "image"},
{"type": "text", "text": "<question text>"}
]
add_generation_prompt=True).AutoProcessor with max_length=2048 and truncation=True.The calibration batch therefore consists of image + text tokenized inputs matching the model's expected multimodal format.
| Attribute | Value |
|---|---|
| GPU | NVIDIA RTX 5090 (Blackwell, SM 120, 32 GB) |
| OS | Windows |
| CUDA | 13.1 runtime driver; PyTorch cu128 wheel |
| Python | 3.10 |
| PyTorch | 2.11.0+cu128 |
| torchvision | 0.26.0+cu128 |
| Key packages | transformers, llmcompressor, datasets, opencv-python, Pillow |
expandable_segments is not supported by the CUDA allocator on Windows, so this option had no effect and was left as a harmless no-op.Key observations from the run:
| Stage | GPU Memory |
|---|---|
| Original BF16 model loaded | ~16.33 GB |
| After AWQ W4A16 applied | ~6.74 GB |
save_compressed=True, producing the compressed-tensors checkpoint layout.| File | Size |
|---|---|
model.safetensors | ~6.8 GB |
config.json | ~7.9 KB |
tokenizer.json | ~11 MB |
chat_template.jinja | ~5.3 KB |
recipe.yaml | ~1.4 KB |
generation_config.json, processor_config.json, tokenizer_config.json | small metadata |
| Total | ~6.8 GB |
This checkpoint is not in legacy AutoAWQ format (quant_method: "awq"). It uses the compressed-tensors format produced directly by llmcompressor:
"quantization_config": {
"quant_method": "compressed-tensors",
"quantization_status": "compressed",
"format": "pack-quantized",
...
}
Compatibility with downstream engines:
transformers can load it as long as compressed-tensors is installed.vLLM support depends on the vLLM version understanding quant_method: "compressed-tensors" for Qwen3-VL; use a recent release and verify before deploying.AWQModifier used in this run is the legacy compatibility shim in llmcompressor. Newer releases recommend replacing it with AWQTransformModifier followed by QuantizationModifier.To reproduce this quantization from the base model:
MBZUAI/VCGBench-Diverse and extract the videos/ directory next to vcgbench_diverse_qa.json.torch==2.11.0+cu128, torchvision==0.26.0+cu128, llmcompressor, transformers, datasets, opencv-python, Pillow).max_seq_length=2048.Same license as the base model jdopensource/JoyAI-VL-Interaction-Preview. Please refer to the base model card for the exact license terms.