Downloads · 30 days
80
100% of all-time downloads
cbert33/Nex-N2.5-mini-FP8-Calibrated
Nex-N2.5-mini-FP8-Calibrated is a image-text-to-text model from cbert33. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains an FP8 quantized derivative of nex-agi/Nex-N2.5-mini, pinned to source revision 87420286149d9cce9bd46cd335ef9bda33c37c1b.
Downloads · 30 days
80
100% of all-time downloads
All-time downloads
80
Public
Parameters
35.1B
36.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors36.6 GB · 100%
How the weights are stored.
F8_E4M333.6B · 96%
From the Hugging Face model README
This repository contains an FP8 quantized derivative of nex-agi/Nex-N2.5-mini, pinned to source revision 87420286149d9cce9bd46cd335ef9bda33c37c1b.
The release uses block-scaled FP8 W8A8 for eligible linear layers and includes calibrated static FP8 KV-cache scales. The multimodal vision tower, hybrid linear-attention state, routing components, embeddings, output head, and other sensitive modules remain in their source precision.
This is a community quantization. Refer to the original Nex-N2.5-mini model card for the model family description, benchmark results, intended applications, and upstream usage guidance.
Qwen3_5MoeForConditionalGenerationqwen3_5_moeThe model was quantized with LLM Compressor using the FP8_BLOCK W8A8 scheme.
Linear weights use static, symmetric 8-bit floating-point quantization.HuggingFaceH4/ultrachat_200k, with a maximum calibration sequence length of 2,048 tokens.The quantization recipe excludes components that are sensitive, unsupported, or intentionally preserved:
Validation found 595 protected tensors that exactly match the source checkpoint.
The finalized artifact passed the following checks:
The vision weights were preserved exactly, but this release has not been given a separate end-to-end image benchmark.
Use a recent vLLM build with support for Qwen3.5 MoE, Compressed Tensors FP8_BLOCK, multimodal input, and FP8 KV cache.
vllm serve cbert33/Nex-N2.5-mini-FP8-Calibrated \
--trust-remote-code \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
Memory use, concurrency, and maximum practical context depend on the serving engine, hardware, KV-cache allocation, and enabled multimodal features.
The upstream model card recommends:
temperature: 0.7top_p: 0.95top_k: 40The included chat template uses reasoning_effort:
none: respond without a reasoning trace;medium: adaptive thinking, the upstream default;high: always enable thinking.Example OpenAI-compatible request:
{
"model": "Nex-N2.5-mini-FP8-Calibrated",
"messages": [
{
"role": "user",
"content": "Explain how binary search works."
}
],
"reasoning_effort": "medium",
"temperature": 0.7,
"top_p": 0.95
}
The request model name must match the name exposed by your serving engine.
The source model uses the qwen3_coder tool-call parser. Enable automatic tool choice and configure that parser in the serving engine. The included chat template should be used unless the runtime supplies a verified equivalent.
The source model is released under the Apache 2.0 license. This derivative retains that license. Nex-N2.5-mini was created by Nex-AGI.