Downloads · 30 days
11
14% of all-time downloads
kylebrodeur/microfactory-node-lora-v3-qat
microfactory-node-lora-v3-qat is a machine learning model from kylebrodeur. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as gemma.
I trained this LoRA on top of the QAT-trained gemma-4-E4B-it-qat-q40-unquantized base. It runs parallel to v2: the same O'Brien judgment, but I wanted to see if fine-tuning on a Quantization-Aware-Trained base keeps m…
Downloads · 30 days
11
14% of all-time downloads
All-time downloads
76
Public
Repo size
67.1 MB
Likes
0
Public
Click a slice to open those files.
.safetensors35 MB · 52%
From the Hugging Face model README
I trained this LoRA on top of the QAT-trained gemma-4-E4B-it-qat-q4_0-unquantized base. It runs parallel to v2: the same O'Brien judgment, but I wanted to see if fine-tuning on a Quantization-Aware-Trained base keeps more quality after q4_0 GGUF conversion.
Give it a print job — material, geometry, room temperature and humidity — and it returns structured Advice JSON:
| Parameter | Value |
|---|---|
| Base model | google/gemma-4-E4B-it-qat-q4_0-unquantized |
| Method | LoRA (PEFT) |
| Rank | r=4, α=8 |
| Epochs | 1 |
| Learning rate | 2e-4 |
| Batch size | 2 × 4 gradient accumulation |
| Max sequence length | 1536 |
| Dataset | 180 train / 80 eval (live-generated on Modal A10G) |
| GPU | NVIDIA A10G (24GB) |
| Framework | TRL SFTTrainer + transformers 5.x |
Same low-rank, single-epoch setup as v2. The variable is the QAT base.
I generated the training set by driving the base model across a grid of 4 materials × 5 geometries × 3 temperatures × 3 humidities (train), with 2 temperatures × 2 humidities held out for eval. Each example is a chat-format pair: system prompt describing the job → structured Advice JSON response.
I kept targets noisy — temperature=0.7, top_p=0.95 — to prevent template memorization.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("google/gemma-4-E4B-it-qat-q4_0-unquantized")
base = AutoModelForCausalLM.from_pretrained(
"google/gemma-4-E4B-it-qat-q4_0-unquantized",
dtype=torch.bfloat16,
device_map="auto"
)
tuned = PeftModel.from_pretrained(base, "kylebrodeur/microfactory-node-lora-v3-qat")
messages = [{"role": "user", "content": "Your prompt here"}]
inputs = tok.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(tuned.device)
out = tuned.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)
print(tok.decode(out[0], skip_special_tokens=True))
This adapter proposes settings. It does not validate them. A deterministic Spine clamps every proposed value against hard material bounds before any printer sees them. The LoRA gives the opinion; the Spine has the veto.
| Version | Base | Rank | Epochs | Dataset | Result |
|---|---|---|---|---|---|
| v1 | gemma-3-1b-it | r=16 | 3 | deterministic | ❌ Parroted template |
| v2 | gemma-4-E4B-it | r=4 | 1 | live-generated | ✅ Well-Tuned |
| v3 | gemma-4-E4B-it-qat-q4_0-unquantized | r=4 | 1 | live-generated | ✅ Well-Tuned (QAT-trained — better fidelity after q4_0 quant) |
v1 taught me what not to do. v3 tests whether QAT pre-training helps the quantized artifact.
This adapter is narrow by design, and it will fail loudly outside that narrow band.
Two quantized GGUFs of this adapter, merged into the QAT base, are published.
Both live in kylebrodeur/microfactory-node-gguf
and on the public Ollama registry:
| Quant | HF Hub file | ollama run … (registry tag) | Why pick this one |
|---|---|---|---|
| q4_k_m | microfactory-node-v3-qat.gguf (5.1 GB) | kylebrodeur/microfactory-node-v3-qat | Balanced default |
| q4_0 (QAT-native) | microfactory-node-v3-qat-q4_0.gguf (4.9 GB) | kylebrodeur/microfactory-node-v3-qat:q4_0 | Highest fidelity — this is the quant the QAT base was trained for |
# Public Ollama registry (one-liner)
ollama run kylebrodeur/microfactory-node-v3-qat # q4_k_m, recommended
ollama run kylebrodeur/microfactory-node-v3-qat:q4_0 # QAT-native quant
# Direct from HF Hub (template/system/params auto-applied)
ollama run hf.co/kylebrodeur/microfactory-node-gguf:microfactory-node-v3-qat.gguf
ollama run hf.co/kylebrodeur/microfactory-node-gguf:microfactory-node-v3-qat-q4_0.gguf
See the
full publishing runbook
for the merge → quantize → upload pipeline. The non-QAT sibling lives at
microfactory-node-lora-v2.
This adapter inherits the Gemma license from its base model.