Downloads · 30 days
29
48% of all-time downloads
posttrainllm/pace-intent-router-v8
pace-intent-router-v8 is a text classification model from posttrainllm. Use it when you need a label for a piece of text. It is set up for posttrainllm. The card lists the license as mit.
A 49.5M-parameter intent classifier trained from scratch on a 127K-example synthetic corpus for Pace's 7-class intent taxonomy. It runs in about 3–4 ms warm. On the earlier source-matched synthetic holdout it beat Qwe…
Downloads · 30 days
29
48% of all-time downloads
All-time downloads
60
Public
Repo size
594 MB
Likes
0
Public
Click a slice to open those files.
.tinygpt594 MB · 100%
From the Hugging Face model README
A 49.5M-parameter intent classifier trained from scratch on a 127K-example synthetic corpus for Pace's 7-class intent taxonomy. It runs in about 3–4 ms warm. On the earlier source-matched synthetic holdout it beat Qwen3-4B-Instruct, but that quality claim does not survive the fresh sealed benchmark below.
This is a posttrainllm-built specialist — trained with the
train-extractor command on a ToolRouterModel medium preset (10 layers,
640 d_model, byte-level vocab 256). No pre-trained weights, no
distillation from a teacher. The model learns Pace-specific decision
boundaries from synthetic data generated by the Pace repository's
scripts/generate-intent-corpus-v2.py.
pace-intent-router-v8runs/pace-intent-router-v8.tinygpt (566 MB)runs/pace-intent-router-v8.tinygpt.labels.json.tinygpt (ToolRouterModel checkpoint)posttrainllm-ToolRouterModel-medium (trained from scratch)train-extractor (cross-entropy classification,
AdamW + cosine decay, 18000 steps on v5 data)The 63-instance, zero-overlap, frontier-qualified Everyday Specialist Benchmark reverses the old ordering:
| Entry | Accuracy | Unknown recall | Mean warm latency |
|---|---|---|---|
Codex gpt-5.5 frontier | 100.0% | 100.0% | 519 ms amortized batch |
| Qwen3-4B-Instruct 4-bit | 93.7% | 77.8% | 211 ms |
| Apple FoundationModels | 92.1% | 77.8% | 522 ms |
| Pace Intent Router v8 | 57.1% | 55.6% | 3.8 ms |
Decision: reject v8 as the production winner on this task. It remains the
latency floor, but the training generator did not cover enough fresh
real-user-like phrasing. See
evals/everyday-benchmark/pace-intent-sealed-v1.md.
The rest of this section is historical source-matched holdout evidence. It is useful for reproducing the factory run, not for an official win claim.
| Metric | ToolRouterModel v8 | Qwen3-4B-Instruct | Delta |
|---|---|---|---|
| Overall accuracy | 95.5% | 84.75% | +10.8 pp |
| Eval size | 14,995 | 997 (stratified sample) | — |
| Latency p50 | 3.1 ms | 240 ms | 77x faster |
| Params | 49.5M | 4B (4-bit) | — |
The router checkpoint is 566 MB on disk; no pinned Qwen artifact-size measurement is committed, so no size-delta claim is made.
| Class | v8 (49.5M) | v5 (49.5M) | Qwen3-4B |
|---|---|---|---|
| chitchat | 93.8% | 95.6% | 78.5% |
| pureKnowledge | 96.4% | 97.0% | 83.0% |
| screenDescription | 97.6% | 97.1% | 95.8% |
| screenAction | 96.8% | 97.4% | 91.0% |
| research | 93.4% | 97.1% | 100.0% |
| phoneLargeModel | 98.4% | 96.4% | 76.7% |
| unknown | 79.6% | 70.3% | 2.6% |
| Overall | 95.5% | 95.9% | 84.75% |
generate-intent-corpus-v2.py (base corpus,
combinatorial expansion) + generate-intent-supplement-v2.py
(targeted weak-spot supplement), both in the sibling Pace repositoryThe synthetic corpus encodes Pace-specific decision boundaries that a general LLM doesn't know:
These boundaries are product-specific, not language-general. A from-scratch model learns them directly from the corpus; a general LLM needs few-shot examples or fine-tuning to learn them.
This model is a router, not a planner. It classifies a user utterance into one of 7 intent classes in 3ms. The intent class determines which pipeline handles the turn:
| Intent | Route | Model |
|---|---|---|
| chitchat | fast path | Apple FM / local text-only |
| pureKnowledge | answer directly | Apple FM / local text-only |
| screenDescription | read screen | local planner + VLM |
| screenAction | execute tool | local planner + action layer |
| research | research tier | codex CLI / Claude CLI |
| phoneLargeModel | cloud bridge | codex CLI / cloud bridge |
| unknown | full pipeline | local planner (best-effort) |
Do not use this model for response generation. It only classifies.
The Pace intent classifier was originally rule-based (hand-written phrase lists). This model was trained to test whether a learned classifier could beat the rules. It can — 95.5% vs the rules' ~95% on the synthetic eval — and it also beats Apple Foundation Models (3B, in-process) on the same task: 95.5% vs 76.5% on a 200-example stratified sample. The model is most useful as:
| Model | Accuracy | Latency | Error rate |
|---|---|---|---|
| TinyGPT v8 (49.5M) | 95.5% | 3.1ms p50 | 0% |
| Qwen3-4B-Instruct (4-bit) | 84.75% | 240ms p50 | 0% |
| Apple FM (3B, in-process) | 76.5% | 1597ms mean | 0.5% |
FM's main weakness: it conflates research with pureKnowledge
(33.3% vs 93.4% — a 60pp gap) because it doesn't understand the
Pace-specific distinction between "research X" (multi-step) and
"what is X" (single answer). FM also refused to answer one query
("model refused to answer").
train-extractor code; no pre-trained or third-party weights are included.generate-intent-corpus-v2.py and
generate-intent-supplement-v2.py); no external datasets or user data.generate-intent-corpus-v2.py,
generate-intent-supplement-v2.py) and the Pace routing architecture
(PaceIntentClassifier.swift) live in the sibling Pace repository, not
this one. The earlier local factory-run receipts were not committed; the
sealed V1 evaluation above is the authoritative public evidence.