Downloads · 30 days
0
mie00/autoeq
autoeq is a text generation model from mie00. Use it when you need the model to write or continue text. It is set up for onnx. The card lists the license as apache-2.0.
A 3.3M-parameter encoder-decoder transformer that turns human-typed math into LaTeX on every keystroke. Trained entirely on synthetic data whose plaintext renderer encodes a human prior over ambiguous notation (/2a me…
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
210 MB
Likes
0
Public
Click a slice to open those files.
.pt157 MB · 90%
From the Hugging Face model README
A 3.3M-parameter encoder-decoder transformer that turns human-typed math into LaTeX on every
keystroke. Trained entirely on synthetic data whose plaintext renderer encodes a human prior over
ambiguous notation (/2a means "divide by 2a", sq means √, x1 means x_1, sin x/2 means
(sin x)/2 …). Beam search (k=3) surfaces alternatives when the input is genuinely ambiguous.
x = -b +- sq(b^2-4ac)/2a → x=\frac{-b\pm\sqrt{b^2-4ac}}{2a}
e^-x^2/2 → e^{-\frac{x^2}{2}}
bold E = -grad phi → \mathbf{E}=-\nabla\phi
n_eff = c/v_p → n_{\mathrm{eff}}=\frac{c}{v_p}
delta x → \delta x (alt: \Delta x)
int_0^2pi sin x dx → \int_0^{2\pi}\sin x\,dx
{x^2 if x>0; 0 otherwise} → \begin{cases}x^2&\text{if}\,x>0\\0&\text{otherwise}\end{cases}
[[a,b],[c,d]] → \begin{bmatrix}a&b\\c&d\end{bmatrix}
\underline{M} → \underline{M} (macros pass through)
| file | what |
|---|---|
autoeq-v1.4.pt | PyTorch checkpoint (config + EMA weights); load with autoeq.model.load_checkpoint |
encoder.onnx, decoder.onnx | fp32 ONNX graphs (opset 18). Encoder also emits the cross-attention K/V; decoder has an explicit KV cache and accepts T_new ≥ 1 tokens |
encoder.int8.onnx, decoder.int8.onnx | int8 dynamic-quantized graphs used by the demo (≈4.2 MB total) |
vocab.json | input character table + normalization map, output token list, decoding-constraint flags |
Graph I/O:
encoder: src_ids[B,S] int32, src_len[B] int32
→ cross_k[L,B,H,S,dh], cross_v[L,B,H,S,dh], src_bias[B,1,1,S]
decoder: tgt_ids[B,T_new] int32, pos_offset[1] int32,
self_k[L,B,H,T_past,dh], self_v, cross_k, cross_v, src_bias
→ logprobs[B,T_new,V], new_self_k[L,B,H,T_past+T_new,dh], new_self_v
Decoding must be constrained beam search over the output tokens (braces balanced, &/\\ only
inside environments, and \ followed only by letters) — a reference implementation lives in the
repo (autoeq/decode.py, and the identical web/decode.js for the browser). Input: characters
(≤128); output: characters + LaTeX macros as single tokens (420 tokens); vocab.json has
everything needed.
Pre-LN transformer, d_model 192, 4 heads, FFN 768, 3 encoder + 3 decoder layers, learned positions, tied decoder embeddings — 3.29M parameters. Trained 60k steps (batch 512, AdamW, label smoothing 0.1, EMA) ≈ 40 minutes on one RTX 4090, on ~30M unique synthetic examples (65% full expressions, 30% typing prefixes with aligned partial LaTeX, 5% famous formulas).
| metric | result |
|---|---|
| normalized exact match, random synthetic formulas (beam-3 / oracle@3) | 91.6% / 98.8% |
| held-out famous formulas (49 formulas never in training) | 100% |
| ambiguity suite (187 hand-written cases): top-1 / intended alternative in top-3 | 99.5% / 100% |
| coverage — a target formula is reachable by some spelling | 98.7% |
| outputs that compile in KaTeX | 100% |
| AsciiMath (rule-based) on the same ambiguity suite | 24% |
Latency: ~10.0 ms/keystroke typing the quadratic formula in onnxruntime-web (WASM, 1 thread), down from 12.6 ms in v1.3 despite the larger vocabulary — the beam search now selects the top k+1 candidates directly instead of sorting the whole vocabulary at every step.
Comparing against v1.3: the evaluation set is regenerated from the generator each release, so these numbers are not comparable across versions. Scored on one evaluation set both models can be run on, v1.4 gets 92.8% greedy / 91.2% beam and v1.3 gets 92.6% / 91.2% — pass-through cost nothing on the notation v1.3 already handled.
Any macro the grammar does not model is now passed through exactly as typed, instead of being
approximated by whatever the vocabulary happened to contain (\underline{M} previously decoded to
\cup\nabla\{M\}).
| typed | LaTeX |
|---|---|
\underline{M}, \widehat{AB} | \underline{M}, \widehat{AB} |
x = \underline{M} + 1 | x=\underline{M}+1 |
\overset{a}{b}, \mathfrak{g}, \mathscr{L} | \overset{a}{b}, \mathfrak{g}, \mathscr{L} |
\aleph, \wp, \vdots, \models, \bigoplus | the same, verbatim |
\varprojlim, \circledast | the same — not vocabulary entries, spelled out via the escape |
164 macros are single output tokens. Everything else is spelled character by character behind a
\ escape token, so coverage is not limited to that list and macros the model never saw in
training still pass through. The plaintext for such a macro is its LaTeX, which is what keeps
one input mapping to one canonical output.
Real-time LaTeX entry in editors, note apps, chat UIs. Multi-line align, \left/\right and
colors remain unreachable, and environments cannot be passed through the way macros can. Prose is
only supported in quotes ("speed of light" → \text{…}). Trained on synthetic data only —
unusual personal notation may be misread; the shown alternatives and explicit parentheses always
let you force a reading. Some distinctions plain text simply cannot carry — bold vs arrow vectors,
upright vs italic labels, delta vs Delta — so those are offered as alternatives rather than
guessed at.
Apache License 2.0 — code and weights.