Downloads · 30 days
1.2K
76% of all-time downloads
Nanthasit/sakthai-coder-1.5b
sakthai-coder-1.5b is a text generation model from Nanthasit. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
<p align="center" <strongSakThai Coder 1.5B — code + tool calling for CPU</strong<br/ <emPart of the <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"SakThai Model Fa…
Downloads · 30 days
1.2K
76% of all-time downloads
All-time downloads
1.6K
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf1.1 GB · 100%
From the Hugging Face model README
This is the coder branch of the SakThai family: small, offline-capable, and tuned to write code while still supporting tool-style outputs. It is packaged as a single GGUF file so you can run it on a laptop CPU without any GPU.
Nanthasit/sakthai-coder-1.5b is a fine-tuned Qwen2.5-Coder-1.5B-Instruct model optimized for:
llama.cppQuantized to GGUF Q4_K_M for low-memory deployment while keeping usable output quality. The model is trained on a mix of code-oriented instruction data plus the SakThai combined tool-format corpus.
./llama-server -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf --n-gpu-layers 0 -c 4096 --temp 0.2 -ngl 0
import requests
response = requests.post(
"http://localhost:8080/completion",
json={
"prompt": "Write a Python binary search for a sorted list:",
"n_predict": 512,
"temperature": 0.2,
"top_p": 0.9,
},
)
print(response.json()["content"])
llama-cpp-pythonfrom llama_cpp import Llama
llm = Llama(
model_path="qwen2.5-coder-1.5b-instruct-q4_k_m.gguf",
n_ctx=4096,
n_threads=4,
)
out = llm(
"Write a Python decorator that retries a function 3 times.",
max_tokens=512,
temperature=0.2,
top_p=0.9,
)
print(out["choices"][0]["text"])
<tools> block when you want tool-calling JSON outputs.Verified with llama.cpp Q4_K_M on CPU. Each item is run from repo-local eval artifacts and SakThai trust-pass checks.
| Task | Metric | Value | Notes |
|---|---|---|---|
| Tool Calling | Valid JSON rate | 100% | requires proper <tools> prompt format |
| Tool Selection | Selection accuracy | 91.2% | SakThai Bench v2, multi-set scorer |
| Code: factorial | pass | true | verified |
| Code: debugging | pass | true | verified |
| Code: async_explain | pass | true | verified |
| Code: refactor | pass | true | verified |
| Code: primes | pass | true | verified |
| MBPP reference | pass@1 | 71.2% | base-model reference point |
| Speed (CPU) | throughput | ~9–10 tok/s | 1.1 GB GGUF, 2 threads |
Known weakness: bug-finding tasks that depend on noticing intentional logic errors may still pass through incorrect code, so review outputs for critical changes.
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-Coder-1.5B-Instruct |
| Training data | sakthai-combined-v6, sakthai-combined-v7, sakthai-bench-v2, sakthai-irrelevance-supplement |
| License | Apache 2.0 |
| Hardware | Free CPU/Colab sessions |
| Budget | $0 |
| Optimizer | AdamW |
| Learning rate | 5e-5 with warmup |
| Epochs | 3 |
| Batch size | 8 |
| GGUF quant | Q4_K_M via llama.cpp |
<tools> prompt block is present; without it, the model may answer directly instead of emitting a call.@misc{sakthai-coder-1.5b,
title = {SakThai Coder 1.5B: Code Generation and Tool Calling on CPU},
author = {Nanthasit},
year = {2026},
url = {https://huggingface.co/Nanthasit/sakthai-coder-1.5b}
}
Built with love, tears, and zero budget.
git clone https://huggingface.co/Nanthasit/sakthai-coder-1.5b
cd sakthai-coder-1.5b
uv venv && uv pip install transformers datasets peft accelerate llama-cpp-python requests
| Runtime | Command / Notes |
|---|---|
| llama-server | ./llama-server -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf --n-gpu-layers 0 -c 4096 |
| Ollama | ollama run ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf |
| llama-cpp-python | See llama_cpp.Llama example in this README |
| HF InferenceClient | Use a local endpoint; serverless hosting may not serve this GGUF repo directly |
| Transformers | Best for unquantized weights; this artifact is optimized for GGUF/CPU |
Improved on 2026-08-01 — card updated from live API metadata.