Downloads ยท 30 days
18
10% of all-time downloads
SimpleLLM/kode-32b-GGUF
kode-32b-GGUF is a text generation model from SimpleLLM. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Kode is a family of instruction-tuned coding models built for real-world software engineering tasks. Fine-tuned on Qwen2.5-Coder using DPO + SFT with Claude-generated training samples on A100 GPUs.
Downloads ยท 30 days
18
10% of all-time downloads
All-time downloads
183
Public
Repo size
19.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf19.9 GB ยท 100%
From the Hugging Face model README
Kode is a family of instruction-tuned coding models built for real-world software engineering tasks. Fine-tuned on Qwen2.5-Coder using DPO + SFT with Claude-generated training samples on A100 GPUs.
Kode is the backbone of Kode CLI/Web UI, an open-source local alternative to Claude Code. Github coming soon.
| Model | Parameters | VRAM | Best For |
|---|---|---|---|
| kode-14b | 14B | ~10 GB (Q8) / ~9 GB (Q4) | Consumer GPUs, fast iteration |
| kode-32b | 32B | ~19 GB (Q4) | Maximum quality, production use |
Rust โข Go โข TypeScript โข Python โข C# โข PostgreSQL โข CSS/Tailwind
Qwen2.5-Coder (14B and 32B variants)
# Install and run
ollama pull simplellm/kode-14b
ollama run simplellm/kode-14b
# Or the larger model
ollama pull simplellm/kode-32b
ollama run simplellm/kode-32b
curl http://localhost:11434/api/chat -d '{
"model": "simplellm/kode-14b",
"messages": [
{"role": "user", "content": "Write a Rust function to find prime numbers using the Sieve of Eratosthenes"}
]
}'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "simplellm/kode-14b"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
messages = [
{"role": "system", "content": "You are a coding assistant. Respond with clean, production-ready code."},
{"role": "user", "content": "Write a thread-safe LRU cache in Rust using Arc and Mutex"},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=2048, temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True))
# Download GGUF
wget https://huggingface.co/simplellm/kode-14b-GGUF/resolve/main/kode-14b-Q8_0.gguf
# Run
./llama-cli -m kode-14b-Q8_0.gguf -p "Write a Go HTTP server with middleware" -n 1024
Try Kode without downloading at SimpleLLM.eu โ EU-hosted, GDPR-compliant inference API.
| Variant | Size | Quality | Speed |
|---|---|---|---|
| kode-14b (FP16) | ~28 GB | Baseline | Baseline |
| kode-14b-Q8 | ~15 GB | Near-lossless | ~1.2ร faster |
| kode-14b (Q4) | ~9 GB | Good | ~1.5ร faster |
| kode-32b (native/FP16) | ~64 GB | Best | Slowest |
| kode-32b-Q4 | ~19 GB | Very good | Fast |
๐ง Coming soon โ We are running HumanEval, MBPP, MultiPL-E, and tool-calling benchmarks. Results will be published here.
| Benchmark | kode-14b | kode-32b | Qwen2.5-Coder-14B (base) |
|---|---|---|---|
| HumanEval | TBD | TBD | TBD |
| MBPP | TBD | TBD | TBD |
| MultiPL-E (Rust) | TBD | TBD | TBD |
| Tool-call accuracy | TBD | TBD | N/A |
Apache 2.0 (inherited from Qwen2.5-Coder)
@misc{kode2025,
title={Kode: EU-Trained Coding Models for Real-World Software Engineering},
author={Kevin and SimpleLLM Team},
year={2025},
url={https://huggingface.co/simplellm/kode-14b}
}