Downloads · 30 days
650
100% of all-time downloads
jaredpalmer/kev-8b
kev-8b is a text classification model from jaredpalmer. Use it when you need a label for a piece of text. It is set up for peft. The card lists the license as apache-2.0.
Previous generation (Qwen3). This checkpoint is kept as the fast option on Apple Silicon (its attention-only backbone runs the packed forward at full speed on MPS). For accuracy and calibration use Kev-9B: on the lock…
Downloads · 30 days
650
100% of all-time downloads
All-time downloads
650
Public
Repo size
378 MB
Likes
13
Public
Click a slice to open those files.
.safetensors175 MB · 88%
From the Hugging Face model README
Previous generation (Qwen3). This checkpoint is kept as the fast option on Apple Silicon (its attention-only backbone runs the packed forward at full speed on MPS). For accuracy and calibration use Kev-9B: on the locked test it scores 0.837 vs 0.780 out of domain against this model on the same items. Weights:
jaredpalmer/kev-8b.
Kev-8B is a decision model: one document (the state) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16) plus a pointer head on Qwen/Qwen3-8B-Base (revision 49e3418f), serving TypeSafe's public /v1/systemone contract.
The most accurate kev. The best checkpoint of any size under a frozen, checksummed protocol: best in-distribution accuracy, best out-of-domain accuracy (0.796 on transfer-v4 dev, six points from Jev), best held-out rule reasoning of any Kev at 8B. Same recipe at two seeds: 0.796 / 0.774; this checkpoint is the seed selected on the development partition.
jaredpalmer/kev-8b (this repo; trial v7-final/00-trial-0)PLAN.md, runs/leaderboard.md| | Kev-0.5B (prototype) | Kev-0.6B | Kev-4B | Kev-8B | Jev | |---|---|---|---|---| | in-distribution accuracy (decision-v4 dev, 1,200 q) | 0.712 | 0.801 | 0.854 | 0.863 | 0.845 | | out-of-domain accuracy (transfer-v4 dev, 560 q) | 0.561 | 0.620 | 0.790 | 0.796 | 0.857 | | out-of-domain Brier | 0.50 | 0.536 | 0.328 | 0.337 | 0.211 | | confident errors out of domain (p ≥ 0.9 and wrong) | – | 10.8% | 8.2% | 9.9% | 3.7% | | held-out policy structures, both siblings correct | – | 0.08 | 0.73 | 0.69 | 0.86 | | option-order flip rate | 0.21 | 0.02 | 0.00 | 0.00 | 0.00 |
Per-source out-of-domain accuracy (Kev-8B / Jev): QNLI 0.91 / 0.93, SciQ 1.00 / 0.99, TweetEval-offensive 0.79 / 0.81, PAWS 0.78 / 0.79, MMLU 0.70 / 0.90, Emotion 0.56 / 0.59, deadline (3-level date arithmetic) 0.60 / 0.93, (A and B) or not C 0.91 / 0.97, if A then not B else C 0.59 / 0.78.
Seeds: two seeds on decision-v7: transfer 0.796 / 0.774, held-out rule pairs 0.69 / 0.64 (Jev 0.86); this checkpoint is seed 0. Trained on decision-v7 (10k public records + 896 policy records over nine template families incl. four ordinal Score threshold families + 1,680 records from 60 random rule structures with negation anywhere); development/test items are byte-identical to v4, so every number here is comparable with earlier checkpoints.
Locked test, read once (runs/locked/kev-8b-v7-preview-ungated/): in-distribution 0.870 (Brier 0.193), out-of-domain 0.780 (Brier 0.327, confident errors 7.6%, held-out pairs 0.62). This partition will not be read again for this checkpoint.
KEV_DTYPE=bf16 (~17 GB) does. Training took ~70 min on one H100.Frozen suite evals/v6/decision-v6 (development/test bytes identical to v4): 13,000 public records (1,000 per source: the ten v4 sources plus ARC-Challenge, OpenBookQA, CommonsenseQA) plus two programmatic policy arms of 448 records, two epochs, LoRA r=16 on attention and MLP projections, pointer head from scratch, cross-entropy on the option distribution, lr 5e-5 (OneCycle), effective batch 8, bf16 autocast with fp32 master weights, gradient checkpointing, one H100 (~70 min). Augmentation: option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records. No Jev outputs were used for training.
Development partitions select models; the locked test partition is read at most once per candidate. Every number carries suite hash, code hashes, and git commit in result.json. See PLAN.md for the corrections we made to our own earlier claims.
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-8b --port 8008 # KEV_DTYPE=bf16 on a 32 GB Mac
Any TypeSafe-compatible client works: TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8008", model="kev-latest").
Apache-2.0 for the adapter and head; Qwen3 base is Apache-2.0; datasets carry their own licenses.