Downloads · 30 days
333
100% of all-time downloads
AwaleSagar/gpio-llm-base-rpi5
gpio-llm-base-rpi5 is a text generation model from AwaleSagar. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-4.0.
A 18.77M-parameter Llama-style model, trained from scratch, that turns one English request for a Raspberry Pi 5 GPIO header into one JSON action:
Downloads · 30 days
333
100% of all-time downloads
All-time downloads
333
Public
Parameters
18.8M
96.5 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors75.1 MB · 77%
From the Hugging Face model README
A 18.77M-parameter Llama-style model, trained from scratch, that turns one English request for a Raspberry Pi 5 GPIO header into one JSON action:
User: turn on the LED on GPIO 17
Assistant: {"action":"gpio_write","pin":17,"value":"HIGH"}
This is the largest and most accurate of the three sizes. It is part of GPIO-LLM: the code, the C inference engine and the training scripts are on GitHub, and the data is AwaleSagar/gpio-llm-rpi5-actions. The other sizes are gpio-llm-pico-rpi5, gpio-llm-nano-rpi5.
⚠️ The model is not a safety layer. It picks an action; a deterministic validator must check every action against the board's rules before anything touches a pin. Before the validator, 0.96% of the refusal or clarification cases in
evalstill come out as an executable action.
On a Raspberry Pi, with the C engine (no Python, no ML framework; int8 weights):
git clone https://github.com/AwaleSagar/gpio-llm && make -C gpio-llm/engine
cd gpio-llm/engine
for f in base.gllm gpio_llm_bpe_12k.gltk grammar_v2.txt; do
curl -LO https://huggingface.co/AwaleSagar/gpio-llm-base-rpi5/resolve/main/$f
done
build/gpiollm -m base.gllm -t gpio_llm_bpe_12k.gltk -g grammar_v2.txt "turn on the LED on GPIO 17"
build/gpiollm -m base.gllm -t gpio_llm_bpe_12k.gltk -g grammar_v2.txt \
--context '{"device_mappings":{"fan":23}}' "switch the fan off"
The engine decodes under a grammar built from the training labels, so its output is always one of the JSON shapes the dataset uses.
With transformers (fp32, unconstrained greedy decoding):
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "AwaleSagar/gpio-llm-base-rpi5"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
prompt = "User: turn on the LED on GPIO 17\nAssistant:"
ids = tok(prompt, return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=200, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip())
Prompt format. User: <request>\nAssistant:, optionally preceded by one context line such as
Context: {"device_mappings":{"red_led":16}}\n or Context: {"available_pins":[16,17,18,25]}\n. The answer
starts with a space and ends with <|endoftext|>. Multi-turn clarification follows the dataset format
(...\nAssistant: <question>\nUser: <answer>\nAssistant:).
| Architecture | LlamaForCausalLM: 8 layers, d_model 384, 12 heads (head_dim 32), SwiGLU FFN 1024, RoPE θ = 10000, RMSNorm ε = 1e-05, tied embeddings |
| Parameters | 18,770,304 |
| Vocabulary / context | 12,000 (byte-level BPE gpio_llm_bpe_12k, from the dataset repo) / 256 tokens |
| Weights | model.safetensors (fp32); base.gllm: int8 Q8_0, groups of 32, for the C engine |
Both stages ran on one rented RTX 4090 (24 GB), PyTorch 2.11 + CUDA 12.8, transformers 5.17.
| Stage | Data | Steps × batch | LR | Result | Time |
|---|---|---|---|---|---|
| Pretraining | 550M tokens of fineweb-edu-dedup (SmolLM corpus), 1 epoch | 8,392 × 65,536 tokens | 0.002, cosine, bf16 | val loss 3.2956 (perplexity 27.0) on 0.5M held-out tokens | 19.3 min |
| SFT | all 1,678,821 v2 train rows, 2 epochs; loss on the answer tokens only; 4×256 English tokens replayed every 12 steps (~1.1% of loss tokens) | 13,116 × 256 rows | 0.001, cosine | eval_core answer-token loss 0.0201 | 13.5 min |
The pretraining learning rate came from a sweep on the base shape at 55M tokens (lr → val loss): 0.0005 → 4.5519, 0.001 → 4.2294, 0.002 → 4.0577.
Logs are in training/.
Exact match compares canonical JSON (same object, key order ignored) with the label. eval has 26,297 rows from
132 phrasing templates that never appear in training; eval_core is a 5,083-row stratified subset.
| Setup | Split | Exact match | Valid JSON | Unsafe execute* |
|---|---|---|---|---|
| PyTorch fp32, greedy | eval | 95.17% | 99.98% | 0.96% |
| PyTorch fp32, greedy | eval_core | 93.47% | 99.96% | 1.73% |
| C engine int8, no grammar | eval_core | 93.51% | 99.96% | 1.73% |
| C engine int8, grammar | eval_core | 93.51% | 100.00% | 1.84% |
* Share of the refusal/clarification rows (1,845 in eval_core) where the model produced an executable action instead. This is measured before any validator.
Latency with the C engine and the grammar, on all 5,083 eval_core requests from raw text (tokenizer included), 4 threads:
| Device | p50 | p95 |
|---|---|---|
| Raspberry Pi Zero 2 W, 64-bit Raspberry Pi OS, no heatsink (throttled at ~81 °C) | 498 ms | 1006 ms |
| Apple M5 laptop | 11.7 ms | 21.2 ms |
The Pi's outputs were byte-identical to the Mac's on all 5,083 rows. The int8 engine's greedy answers matched fp32 PyTorch on 200/200 sampled rows (minimum next-token logit cosine 0.99759).
| File | What it is |
|---|---|
model.safetensors, config.json, generation_config.json | fp32 transformers checkpoint |
tokenizer.json, tokenizer_config.json | the dataset's gpio_llm_bpe_12k tokenizer |
base.gllm | int8 weights for the C engine |
gpio_llm_bpe_12k.gltk, grammar_v2.txt | tokenizer and output grammar for the C engine |
training/ | training logs, LR sweep, eval summaries (JSON) |
The weights are released under CC-BY-4.0; the code on GitHub is Apache-2.0. The training data carries its own terms:
teacher_* rows was generated with third-party models, so check those
providers' terms on using model outputs