Downloads · 30 days
26
96% of all-time downloads
SMLBuilder/cobol-jeeves-sml
cobol-jeeves-sml is a text generation model from SMLBuilder. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
⚠️ PROTOTYPE / research artifact — read the Honesty section before citing any number. This is a method demonstration, not a working COBOL author and not a generalization result. It was adversarially reviewed (Claude,…
Downloads · 30 days
26
96% of all-time downloads
All-time downloads
27
Public
Parameters
14.3M
57.3 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors57.3 MB · 100%
From the Hugging Face model README
⚠️ PROTOTYPE / research artifact — read the Honesty section before citing any number. This is a method demonstration, not a working COBOL author and not a generalization result. It was adversarially reviewed (Claude, Codex, GLM) and the framing below is the corrected, honest version after that review.
A 14-million-parameter, from-scratch GnuCOBOL model (MLX MicroBrain, ~55 MB)
in the SML — Smallest Language Model family (TinkyBrain lineage). No internet
data; word/symbol COBOL tokenizer where the vocab is the domain boundary.
The model cannot write compiling COBOL on its own (0% single-shot). The experiment is about the calling strategy around a tiny component:
| Strategy (identical 14M weights) | metric | value |
|---|---|---|
| 1 greedy shot | single-output compile | 0% |
| greedy + symbolic repair | single-output compile | 6.2% |
| 8 branches × {raw, repaired}, compiler picks first that builds | pass@16 (compiler-verified, best-of-16) | 33.8% (22/65) |
The 33.8% is pass@16 with a compiler oracle and symbolic repair — NOT a compile-rate in the usual (pass@1) sense. It is the rate at which the search procedure finds a compilable program among up to 16 candidates.
The headline number is not evidence the model generalizes, and not comparable to a larger model:
What it can legitimately claim: a branch + symbolic-repair + compiler-verify call raises an ~0% tiny model to 33.8% pass@16 on an in-distribution set. Whether 14M × 16 verified calls is a compute-competitive path vs one large-model call is an interesting open question this artifact does not settle.
To make any of the numbers mean more, the next step is a template-disjoint, deduplicated held-out set and a matched-protocol comparison (same branches + repair for every model).
MicroBrain (MLX decoder-only): d_model 512 · 8 heads · 6 layers · d_ff 1024 · max_seq 512 · vocab 1431, ~14.3M params, greedy decode. Trained from scratch on
1,236 prompt→COBOL pairs (free-format, cobc -free -c as ground truth).
model.safetensors, config.json, tokenizer.json — the componentsml_call.py — the branch/repair/verify call (the studied variable)cobol_repair.py — mechanical repair (never invents logic)cobol_sml_mcp.py — exposes it as an MCP tool (cobol_draft); honest compiles flagRequires MLX (Apple Silicon) and GnuCOBOL 3.x (cobc) for the verify step.
Apache-2.0. From-scratch weights; deterministic training data. SML / TinkyBrain family.
📄 Verified Program Synthesis with a Symbolic Knowledge Graph and an Un-gameable Compiler Oracle — included here as PAPER.pdf, and published at https://perslis.com/cobol-paper.html. Reports all rates with Wilson 95% confidence intervals, a formal soundness lemma for the verification gate, measured cost, and prominent limitations. Preprint / working draft.