Downloads · 30 days
0
Lei285714/mod-arith-k2
mod-arith-k2 is a machine learning model from Lei285714. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A ~471K-parameter recurrent cell run in a fixed bit-serial Horner loop, computing (a·b) mod p for primes to 2^2048 and operands to 4096 bits. Fork of Robby Sneiderman's neural-horner v8 (MIT), which is submitted separ…
Downloads · 30 days
0
Access
Public
Updated Jul 26, 2026
Repo size
6.2 MB
Likes
1
Public
Click a slice to open those files.
.pt6.2 MB · 100%
From the Hugging Face model README
A ~471K-parameter recurrent cell run in a fixed bit-serial Horner loop, computing (a·b) mod p for primes to 2^2048 and operands to 4096 bits. Fork of Robby Sneiderman's neural-horner v8 (MIT), which is submitted separately as an attributed baseline.
Public benchmark: overall 98.9%, highest tier ≥90% = 10, 129.7 s on a 4090 against a 300 s budget.
One shared p-conditioned cell (bidirectional 3-layer GRU, d_model 96, hidden 128, radix-4 digit and op embeddings, per-bit head) learns s' = (4s + d·x) mod p, d ∈ {0,1,2,3}; a learned op-embedding selects its three uses: reduce a, reduce b, multiply the residues. Radix-4 halves the step count at equal parameter count. Warm-started from v8 with the op-embedding zero-initialised, so at k=1 it reproduces v8 bit-for-bit.
Compliance: no controller, the schedule takes no feedback from the model, every arithmetic step including the conditional subtraction comes from the trained cell, and both routing choices key on the prime's bit-length alone — the shape ruled acceptable in the Zulip fixed-loop-encoder thread.
weights.pt is a weight-space average (model soup) of an annealed checkpoint and a same-recipe re-anneal; it beat both endpoints on identical problems (McNemar p = 0.02, n = 1008). weights_t1.pt serves primes of bit-length ≤ 3. Width policy: groups of ≥257 bits run at min(2080, pbits+32), 17–256 at pbits, <17 at 32.
1. Width robustness is a property of the training width distribution, not of the loop. Sizing the state register produced this project's largest gain, with no retraining. Sweeping Δ = state width − prime bit-length on identical problems, 576 per width, four tier bands:
| Δ | 0 | 1 | 2 | 3 | 4 | 8 | 16 / 24 / 32 / 48 / 64 |
|---|---|---|---|---|---|---|---|
| errors | 27 | 516 | 470 | 116 | 13 | 13 | 3 |
Δ ∈ {1,2} falls to 2–30% accuracy. Δ = 3 is usable but degrades with rollout depth (T7 91.7%, deepest T10 47.9%). Error sets at Δ ∈ {16,24,32,48,64} are identical (Jaccard 1.00), so the 32 we deploy is about twice what is needed.
The same problems on v8 give zero errors at every Δ ∈ {0,1,2,3,4,32} — a flat curve. Loop, inference code and problem set are identical; the difference is training. v8's one-step training holds the state width at L and stratifies primes by bit-length, covering every Δ. Our radix-4 retraining trained each prime size at its own field-filling width, covering only Δ = 0.
Practical form: tier primes cluster just under multiples of 32, so rounding the state width up to a machine-word multiple lands them at 1–2 bits of headroom, the worst region of the curve. Add a fixed margin instead — "round up to 32" is not the same policy as "add 32", and falsifying the former cost us a week. The upstream argument that dynamic-L is exactness-preserving holds for v8 but does not travel with the weights.
2. Small primes are unreachable in any sampling stream. random_prime-style constructors emit odd values with the top bit set, so p = 2 must be injected explicitly. Our T1 sat at 61%, and our own battery could not see p = 2 either, so the hole was first misfiled as a false alarm. The fused single-step universe for p ∈ {2,3,5,7} is 348 transitions — enumerate it.
Continued training failed in three forms. Unanchored tiny-prime fine-tuning cured small primes and ate large-scale precision (T8 −8, T10 −13). L2-SP has no λ window: λ = 10/30/100 relocates the interference rather than removing it (512-bit cells 166→170→172 while 64/128/256-bit cells fall monotonically). A same-recipe re-anneal paid nothing. The mechanism is measurable: our single-step validation floor is ~8e-6 while effective ε is already 1.8e-6, so objective and checkpoint selection both sit below the instrument. Coverage and precision are zero-sum here. Ensembling failed too, in both the width-committee and cross-checkpoint-voting forms.
Structured operands are a live residual, independent of the width effect: the Fermat family is clean (96/96) but the wider power-of-two neighbourhood (2^m, 2^m±1) is 280/288, failing only at primes ≥1024 bits, and headroom does not remove it (at 2013 bits, Δ = 0/4/8/32 gives 8/2/3/3 over 72 problems). Official tiers cannot generate such operands, so this is robustness, not score.
On the probe set above, v8 outperforms this model at every width: the radix-4 line bought a 2× reduction in step count, not accuracy. The deep-chain ε tail is deferred, not solved (T10 min 92%). Small primes needed a second routed checkpoint because one weight set could not hold both regimes. The remaining structural lever is raising the radix further; we ran out of time.
Derived from neural-horner (github.com/Robby955/neural-horner) by Robby Sneiderman, MIT licensed. The architecture and the compliance framing originate there; the radix-4 fusion, width policy, tiny-prime handling and the measurements above are ours.