Downloads · 30 days
390
100% of all-time downloads
quant-mind/Nex-N2.5-mini-W4A16-AutoRound
Nex-N2.5-mini-W4A16-AutoRound is a text generation model from quant-mind. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Quantização de máxima precisão em W4A16 (INT4) do modelo multimodal e agentic nex-agi/Nex-N2.5-mini usando a biblioteca Intel AutoRound.
Downloads · 30 days
390
100% of all-time downloads
All-time downloads
390
Public
Parameters
6B
20.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors20.6 GB · 100%
How the weights are stored.
I3233.6B · 96%
From the Hugging Face model README
Quantização de máxima precisão em W4A16 (INT4) do modelo multimodal e agentic nex-agi/Nex-N2.5-mini usando a biblioteca Intel AutoRound.
qwen3_5_moe (Qwen3_5MoeForConditionalGeneration)shared_expert) por blocosym=True, otimizada para o kernel Marlin no vLLM e SGLang)iters): 200 (otimização de erro quadrático mínimo)enable_minmax_tuning=True e enable_norm_bias_tuning=False (preserva RMSNorms e Gated Norms em BF16)NeelNanda/pile-10k)auto_gptq (compatível com GPTQ / Marlin kernels)shared_expert e shared_expert_gate (ativos em todos os tokens, preservando raciocínio central)mlp.gate (router dos especialistas, preservado para estabilidade do vLLM)mtp (Multi-Token Prediction)lm_head e embeddingsvisual encoder)vllm serve quant-mind/Nex-N2.5-mini-W4A16-AutoRound \
--port 8000 \
--max-model-len 4096 \
--reasoning-parser qwen3 \
--gpu-memory-utilization 0.90
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "quant-mind/Nex-N2.5-mini-W4A16-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype=torch.bfloat16,
trust_remote_code=True
)