Downloads · 30 days
34
24% of all-time downloads
Cadododoom/agents-a1-modelopt-nvfp4
agents-a1-modelopt-nvfp4 is a machine learning model from Cadododoom. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains an optimized NVFP4 (Group Size 16) quantized checkpoint of nvidia/Agents-A1-FP4 (a 35B parameter hybrid Mamba-Attention Mixture-of-Experts agentic model) produced using NVIDIA ModelOpt and for…
Downloads · 30 days
34
24% of all-time downloads
All-time downloads
142
Public
Parameters
17.9B
21.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors21.1 GB · 100%
How the weights are stored.
U816.8B · 94%
From the Hugging Face model README
This repository contains an optimized NVFP4 (Group Size 16) quantized checkpoint of nvidia/Agents-A1-FP4 (a 35B parameter hybrid Mamba-Attention Mixture-of-Experts agentic model) produced using NVIDIA ModelOpt and formatted for native vLLM serving.
Qwen3_5MoeForCausalLM / hybrid Mamba-Attention)visual.*), vocabulary intact, MTP heads stripped (mtp.*)Critical Serving Note: While Group Size 128 (GS128) shrinks footprint further, the vLLM Marlin FP4 CUDA kernel (
marlin_mm) only supports a group size of 16. Attempts to serve GS128 will result in a engine crash (Invalid thread config). Therefore, GS16 is the only viable serving configuration.
vllm serve Cadododoom/agents-a1-modelopt-nvfp4 \
--served-model-name nvidia/Agents-A1-FP4 \
--tensor-parallel-size 2 \
--quantization compressed-tensors \
--moe-backend marlin \
--attention-backend flashinfer \
--kv-cache-dtype fp8 \
--max-model-len 112000 \
--max-num-seqs 4 \
--max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.96 \
--enable-prefix-caching \
--trust-remote-code \
--host 0.0.0.0 \
--port 30000