Downloads · 30 days
192
42% of all-time downloads
nectec/Pathumma-llm-text-4.0.0
Pathumma-llm-text-4.0.0 is a text generation model from nectec. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A Thai reasoning model from the ThaiLLM national initiative. It emits an explicit thinking trace before its final answer, targeting mathematical reasoning, instruction following, and structured tool use in Thai and En…
Downloads · 30 days
192
42% of all-time downloads
All-time downloads
452
Public
Parameters
4.5B
9.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors9.1 GB · 100%
From the Hugging Face model README
A Thai reasoning model from the ThaiLLM national initiative. It emits an explicit thinking trace before its final answer, targeting mathematical reasoning, instruction following, and structured tool use in Thai and English.
Pathumma-llm-4b-think-4.0.0 has the following features:
Use a recent version of transformers; older versions will fail to load the model architecture.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "nectec/pathumma-llm-4b-think-4.0.0"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
)
prompt = "ทำไมวงกลมถึงมี 360 องศา"
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=32768,
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
# Split the reasoning trace from the final answer.
think_end_id = tokenizer.convert_tokens_to_ids("</think>")
try:
index = len(output_ids) - output_ids[::-1].index(think_end_id)
except ValueError:
index = 0
thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")
print("thinking content:", thinking_content) # no opening <think> tag
print("content:", content)
Avoid greedy decoding, which can cause repetition loops in the reasoning trace. Reasoning traces also run long, so capping max_new_tokens too low truncates the answer mid-thought.
vllm serve nectec/pathumma-llm-4b-think-4.0.0 \
--served-model-name pathumma-llm-4b-think-4.0.0 \
--host 0.0.0.0 \
--tensor-parallel-size <TP_SIZE> \
--max-model-len 262144 \
--gpu-memory-utilization 0.85 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml
For local use, Ollama, LM Studio, and llama.cpp are supported once GGUF conversions are available.
Evaluated on Thai-adapted benchmarks covering mathematical reasoning, instruction following, commonsense reasoning, and language consistency.
| Benchmark | Metric | Score |
|---|---|---|
| AIME 2024 (TH) | avg@k | 56.67 |
| MATH-500 (TH) | pass@1 | 85.00 |
| IFEval (TH) — prompt, strict | accuracy | 58.60 |
| IFEval (TH) — prompt, loose | accuracy | 63.72 |
| IFEval (TH) — instruction, strict | accuracy | 67.75 |
| IFEval (TH) — instruction, loose | accuracy | 71.71 |
| HellaSwag (TH) | accuracy | 52.44 |
| Code Switching | — | 97.86 |
<sub>All scores are percentages; higher is better.</sub>
Post-training starts from the ThaiLLM continual-pre-trained base model and proceeds in two stages.
| Subset | Examples | Share |
|---|---|---|
| Instruction Following | 3,501,609 | 49.0% |
| Reasoning (English) | 3,018,230 | 42.3% |
| Tool Use | 334,249 | 4.7% |
| Reasoning (Thai) | 286,747 | 4.0% |
| Total | 7,140,835 | 100% |
Reasoning supervision is drawn mainly from English corpora. Thai capability comes primarily from the continual pre-training carried out in the base model, reinforced here by the Thai reasoning subset and by cross-lingual transfer.
Direct Preference Optimization on 6,303 preference pairs, targeting response formatting and style consistency rather than broad behavioural alignment.
The specific datasets used in post-training are proprietary. The example counts above represent the training data used in each stage.
Post-training was conducted on the LANTA high-performance computing cluster using 16 nodes (64 × NVIDIA A100 40GB) for distributed training.
Released under Apache-2.0, inherited from the base model. Proprietary training data is not distributed with this release.
@misc{pathumma_llm_4b_think_400,
title = {Pathumma-LLM-4B-Think-4.0.0},
author = {NECTEC LLM Team},
year = {2026},
url = {https://huggingface.co/nectec/pathumma-llm-4b-think-4.0.0}
}
Pathumma-llm-4b-think-4.0.0 is part of ongoing research toward sovereign Thai large language models optimized for analytical and tool-augmented intelligence.
LLM Team <br> Jirat Arayapityak ([email protected])<br> Kittitat Manokun ([email protected])<br> Supanat Tangkitvutikul ([email protected])<br> Chanut Sunatho ([email protected])<br> Arnon Saeoung ([email protected])<br> Chaianun Damrongrat ([email protected])<br> Sarawoot Kongyoung ([email protected])
