Downloads · 30 days
19
14% of all-time downloads
TevunahAi/NextCoder-32B-FP8
NextCoder-32B-FP8 is a text generation model from TevunahAi. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
High-quality FP8 quantization of Microsoft's NextCoder-32B, optimized for production inference
Downloads · 30 days
19
14% of all-time downloads
All-time downloads
136
Public
Parameters
32.8B
37.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors37.4 GB · 100%
How the weights are stored.
F8_E4M331.2B · 95%
From the Hugging Face model README
High-quality FP8 quantization of Microsoft's NextCoder-32B, optimized for production inference
This is an FP8 (E4M3) quantized version of microsoft/NextCoder-32B using compressed_tensors format. Quantized by TevunahAi on enterprise-grade hardware with 2048 calibration samples.
For 32B models, vLLM is essential for practical deployment. FP8 quantization makes this flagship model accessible on high-end consumer GPUs.
pip install vllm
Python API:
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
# vLLM auto-detects FP8 from model config
llm = LLM(model="TevunahAi/NextCoder-32B-FP8", dtype="auto")
# Prepare prompt with chat template
tokenizer = AutoTokenizer.from_pretrained("TevunahAi/NextCoder-32B-FP8")
messages = [{"role": "user", "content": "Write a Python function to calculate fibonacci numbers"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
# Generate
outputs = llm.generate(prompt, SamplingParams(temperature=0.7, max_tokens=512))
print(outputs[0].outputs[0].text)
OpenAI-Compatible API Server:
vllm serve TevunahAi/NextCoder-32B-FP8 \
--dtype auto \
--max-model-len 4096
Then use with OpenAI client:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="token-abc123", # dummy key
)
response = client.chat.completions.create(
model="TevunahAi/NextCoder-32B-FP8",
messages=[
{"role": "user", "content": "Write a Python function to calculate fibonacci numbers"}
],
temperature=0.7,
max_tokens=512,
)
print(response.choices[0].message.content)
At 32B parameters, transformers will decompress to ~64GB+ VRAM, requiring multi-GPU setups or data center GPUs. This is not recommended for deployment.
<details> <summary>Transformers Example (Multi-GPU Required - Click to expand)</summary>from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# Requires multi-GPU or 80GB+ single GPU
model = AutoModelForCausalLM.from_pretrained(
"TevunahAi/NextCoder-32B-FP8",
device_map="auto", # Will distribute across GPUs
torch_dtype="auto",
low_cpu_mem_usage=True,
)
tokenizer = AutoTokenizer.from_pretrained("TevunahAi/NextCoder-32B-FP8")
# Generate code
messages = [{"role": "user", "content": "Write a Python function to calculate fibonacci numbers"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
do_sample=True
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Requirements:
pip install torch>=2.1.0 transformers>=4.40.0 accelerate compressed-tensors
System Requirements:
⚠️ Critical: Use vLLM instead. Transformers is only viable for research/testing with multi-GPU setups.
</details>| Property | Value |
|---|---|
| Base Model | microsoft/NextCoder-32B |
| Quantization Method | FP8 E4M3 weight-only |
| Framework | llm-compressor + compressed_tensors |
| Storage Size | ~32GB (sharded safetensors) |
| VRAM (vLLM) | ~32GB |
| VRAM (Transformers) | ~64GB+ (decompressed to BF16) |
| Target Hardware | NVIDIA H100, A100 80GB, RTX 6000 Ada |
| Quantization Date | November 23, 2025 |
| Quantization Time | 213.8 minutes |
Professional hardware ensures consistent, high-quality quantization:
FP8 quantization transforms 32B from "data center only" to "high-end workstation deployable".
This model is sharded into multiple safetensors files (all required for inference). The compressed format enables efficient storage and faster downloads.
The 32B model represents the flagship tier:
| Model | VRAM (vLLM) | Quality | Use Case |
|---|---|---|---|
| 7B-FP8 | ~7GB | Good | General coding, fast iteration |
| 14B-FP8 | ~14GB | Better | Complex tasks, better reasoning |
| 32B-FP8 | ~32GB | Best | Flagship performance, production |
32B Benefits:
This quantization is based on microsoft/NextCoder-32B by Microsoft.
For comprehensive information about:
Please refer to the original model card.
This model inherits the MIT License from the original NextCoder-32B model.
If you use this model, please cite the original NextCoder work:
@misc{nextcoder2024,
title={NextCoder: Next-Generation Code LLM},
author={Microsoft},
year={2024},
url={https://huggingface.co/microsoft/NextCoder-32B}
}
Professional AI Model Quantization by TevunahAi
Making flagship models accessible through enterprise-grade quantization
</div>