Downloads · 30 days
182
4% of all-time downloads
froogai/NousCoder-14B-AWQ
NousCoder-14B-AWQ is a text generation model from froogai. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
NousCoder-14B-AWQ is a 4-bit AWQ (Activation-aware Weight Quantization) quantized version of NousResearch/NousCoder-14B.
Downloads · 30 days
182
4% of all-time downloads
All-time downloads
4.7K
Public
Parameters
14.8B
10 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors10 GB · 100%
How the weights are stored.
I3213.2B · 89%
From the Hugging Face model README
NousCoder-14B-AWQ is a 4-bit AWQ (Activation-aware Weight Quantization) quantized version of NousResearch/NousCoder-14B.
This model specializes in competitive programming and coding tasks, achieving 67.87% Pass@1 on LiveCodeBench v6. It has been post-trained on Qwen3-14B using reinforcement learning on 24k verifiable coding problems.
| Metric | Value |
|---|---|
| Base Model | Qwen3-14B |
| Quantization | 4-bit AWQ |
| Size | 9.4 GB (from 28 GB) |
| VRAM | ~6GB per GPU (2x GPUs) |
| Context Length | 16,384 tokens |
| LiveCodeBench v6 Pass@1 | 67.87% |
| Training | 24k coding problems (RL) |
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model = AutoAWQForCausalLM.from_quantized(
"froogai/NousCoder-14B-AWQ",
device_map="auto",
safetensors=True,
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
"froogai/NousCoder-14B-AWQ",
trust_remote_code=True,
)
# Generate code
prompt = "Write a Python function to implement binary search:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.2,
top_p=0.95,
)
code = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print(code)
python -m vllm.entrypoints.openai.api_server \
--model froogai/NousCoder-14B-AWQ \
--quantization awq_marlin \
--tensor-parallel-size 2 \
--max-model-len 16384 \
--gpu-memory-utilization 0.85 \
--trust-remote-code
# Start vLLM server
vllm serve froogai/NousCoder-14B-AWQ \
--quantization awq_marlin \
--tensor-parallel-size 2
# Make API requests
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "froogai/NousCoder-14B-AWQ",
"prompt": "def quicksort(arr):",
"max_tokens": 512
}'
This model was quantized using AutoAWQ with the following configuration:
{
"zero_point": True,
"q_group_size": 128,
"w_bit": 4,
"version": "GEMM",
}
| Benchmark | Score |
|---|---|
| LiveCodeBench v6 Pass@1 | 67.87% |
| Base Model Pass@1 | 60.79% |
| Improvement | +7.08% |
| Hardware | Speed (tokens/sec) |
|---|---|
| 2x RTX 5060 Ti (awq_marlin) | 15-25 |
| 2x RTX 5060 Ti (awq) | 8-12 |
| Single A100 (awq_marlin) | 40-60 |
| Configuration | VRAM Usage |
|---|---|
| 2x RTX 5060 Ti (TP=2) | ~6GB per GPU |
| Single RTX 5060 Ti | ~12GB |
| Single A100 | ~6GB |
This model excels at:
For coding tasks, use these settings:
If you use this model, please cite:
@misc{nouscoder_14b_2025,
title={NousCoder-14B: Competitive Programming AI Model},
author={Li, Joe},
organization={NousResearch},
year={2025},
month={January},
url={https://huggingface.co/NousResearch/NousCoder-14B}
}
This model is licensed under the Apache 2.0 License. See the LICENSE file for details.
Quantized by: froogai
For questions or issues, please:
Note: This is a quantized version of the original model. For best performance, use the vLLM inference engine with the awq_marlin quantization backend.