Downloads · 30 days
9
16% of all-time downloads
prasadvittaldev/dogmatix
dogmatix is a machine learning model from prasadvittaldev. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as gemma.
A QLoRA fine-tune of google/gemma-4-E2B-it (~2.3B effective) into a local, Python-specialised web-application coding agent. Trained on executed, test-verified trajectories distilled from a GLM-5.2 teacher running live…
Downloads · 30 days
9
16% of all-time downloads
All-time downloads
56
Public
Parameters
5.1B
20.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors10.2 GB · 100%
From the Hugging Face model README
A QLoRA fine-tune of google/gemma-4-E2B-it (~2.3B effective) into a local, Python-specialised web-application coding agent. Trained on executed, test-verified trajectories distilled from a GLM-5.2 teacher running live inside the real harness. Runs on a single 16GB GPU.
v0.2 supersedes v0.1 (still available on the
v0.1branch). v0.2 is a sharper web-app specialisation: better at end-to-end app-building — the thing this project is for — at a measured cost to Django and general Python. The numbers below are honest, single-harness, and include the trade.
Small, well-specified Python web work — FastAPI, Flask, SQLAlchemy/SQLModel — driven through an agent loop: read files, make the edit, run the tests, fix what broke. Fast, local, private; not a frontier engineering model.
All figures are measured under one harness (the model sees the full 7-tool surface — read/write/edit/grep/bash/list/run_tests — with turn-based step counting and a blind-call guard). v0.1's original card reported a web-app number taken under an earlier, inflated harness; it is corrected here.
| benchmark | untuned base | v0.1 | v0.2 |
|---|---|---|---|
| End-to-end app-building pass@1 (12 specs) | 0.182 | 0.273 | 0.364 |
| Web-app agentic pass@1 (156 tasks) | 0.756 | 0.788 | 0.797 |
| Django pass@1 (60 tasks) | 0.567 | 0.550 | 0.500 |
| MBPP (general Python, 100) | 0.760 | 0.710 | 0.700 |
The headline is the first line. End-to-end app-building — build a working app from a natural-language spec, graded by a hidden acceptance suite the model never sees — is what this model is for, and it is the one benchmark with real headroom (base only 0.182). The tuned checkpoints improve monotonically, and v0.2 doubles the base.
Honest caveat: that benchmark is small (11–12 graded specs, one attempt each), so the gain is not yet statistically significant (base-vs-v0.2 McNemar p ≈ 0.63). A larger multi-attempt sweep is running to settle it. Treat "doubles base" as a strong signal, not a proven number. On the older web-app suite v0.2 is at parity with base (+0.04, not significant) — that suite is saturated (base 0.756), which is why the app-building benchmark was built.
Every number above is from the eval loop; the thing users run is the serving proxy. Running the 12 app-build specs through the actual proxies on v0.2:
| path | app-building pass@1 |
|---|---|
| eval-flat reference | 0.333 |
| Pi (OpenAI proxy) | 0.417 |
| Claude Code (Anthropic proxy) | 0.500 |
Both proxy paths are at or above the eval reference — the serving and tool- name translation layer loses no capability. Tool-calling works end-to-end on both harnesses, including WebSearch pass-through (v0.2 reaches for and uses a WebSearch tool when offered, though it was not explicitly trained for search).
v0.1 branch (Django 0.550) or the base.Ships with an OpenAI-compatible and an Anthropic-compatible proxy
(dogmatix/serving/). The translation layer is part of the model: it enforces a
turn boundary the base chat template does not provide (without which the model
emits long bursts of tool calls and never observes their results), guards against
acting on unobserved tool calls, and maps client tool names (Claude Code / Pi)
onto the trained tool vocabulary — passing through tools it has no trained
equivalent for (Bash, WebSearch, …).
QLoRA (4-bit NF4, r=32) on the language-model projections only, loss masked to assistant tokens, merged to bf16 on CPU. Single adapter, single run, curriculum expressed as data-sampling weights. Trajectories are executed and test-verified — never free-generated.
python scripts/appbuild_bench.py (12 specs with hidden
acceptance suites verified against reference implementations before any run).python scripts/appbuild_proxy_bench.py.python scripts/benchmark.py (Django via
--suite benchmark/tasks-django).Prasad Vittaldev