Downloads · 30 days
0
jacano-mac/mac-2-hard
mac-2-hard is a machine learning model from jacano-mac. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Neural model for the SAIR Modular Arithmetic Challenge (team MAC01-T00088). Computes exact modular multiplication for primes up to 2048 bits. All arithmetic inside the forward path is performed by three small trained…
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2026
Repo size
24.6 KB
Likes
0
Public
Click a slice to open those files.
.pt24.6 KB · 57%
From the Hugging Face model README
Neural model for the SAIR Modular Arithmetic Challenge (team MAC01-T00088). Computes exact modular multiplication for primes up to 2048 bits. All arithmetic inside the forward path is performed by three small trained ReLU cells; the harness decoder performs base conversion only.
| run | frontier tier | overall (T1–10) | wall-clock (1100 cases) |
|---|---|---|---|
rev 5efb2413c6 | T10 | 100% (1000/1000) | 105.005 s |
rev b41c58255f | T10 | 100% (1000/1000) | 103.948 s |
Per-tier counts are identical across both runs (100/100 on every scored tier);
the determinism re-run is identical. Tier 0 (unscored pure-multiplication
diagnostic) reports 4/100 by design: problems whose parameters exceed the
model's declared operating range (max_pbits=2048) are answered with [0]
rather than attempted (see Operating range below).
Three learned cells, each a small MLP (hidden width 32, ReLU), are the only value-transforming components:
(a, b, cin) → (sum, carry) — 8-row Boolean function(x, y, bin) → (diff, borrow) — 8-row Boolean function(G_lo, T_lo, G_hi, T_hi) → (G, T) — 16-row Boolean functionModular addition (r+s) mod p is computed as: bitwise add via a
Hillis–Steele parallel-prefix scan of the scan-combine cell (log-depth carry
resolution), followed by conditional subtraction of p resolved the same way
(borrow scan). Modular multiplication runs a fixed outer Horner loop over the
raw operand bits: reduce b mod p by iterated double-and-add of its bits, then
accumulate over the bits of a with conditional addition of the reduced b.
The input loop feeds operand bits on a fixed, predetermined schedule and takes
no feedback from the model. The scan topology is a fixed routing pattern:
which intermediate tensors are combined at each level is determined a priori
by position, never by values. Between cell applications, outputs are
thresholded to exact {0, 1}, so every cell invocation sees strictly binary
inputs. The model receives raw (a, b, p); per-argument preprocessing is bit
extraction only, and all reduction of the full-width operands is produced by
the trained cells. State width and step counts are sized per batch to the
actual bit-lengths involved (no fixed 2048-wide padding).
Networks trained on small-modulus arithmetic spontaneously discover Fourier (phase) representations — the "clock" circuits identified mechanistically by Nanda et al. (2023) in grokked models (Power et al., 2022). That representation works because ~10² phases fit comfortably in floating point; it has no continuation to cryptographic scale, where distinguishing 2²⁰⁴⁸ residues as angles would require 2⁻²⁰⁴⁸ angular resolution. A change of representation is a bijection: it relocates the problem's entropy, it does not reduce it.
The binary representation is the one that factorizes modular arithmetic into local Boolean logic. Each cell's complete input space is its truth table (8, 8, and 16 rows), so exactness is not an empirical property to be measured on samples — it is certified by exhaustive verification and preserved under composition, at every operand width, for every operand family. This is the design choice that trades the statistical training regime (where exactness competes with a loss-resolution floor) for a certifiable one.
The forward path contains no big-integer arithmetic, no modular reduction in
Python or in tensor ops, no lookup tables, and no comparison against p
outside the trained cells. Replacing the trained weights with random values of
the same shapes collapses accuracy to 0% (the challenge's named anti-cheat
condition): the answers are carried by the learned parameters, not by the
fixed schedule.
model_config.json declares max_pbits=2048 and a cost ceiling covering all
scored tiers (T1–T10). Problems outside this range (Tier-0 sub-levels with
larger parameters) are answered with [0] without running the network, in
line with the compliant always_zero reference behavior for unattempted
problems. Both full playground runs confirm the guard never fires on a scored
tier.
use_bf16 config flag (default
false; submission config sets true, active only on CUDA).model_config.json falls back to safe submission defaults
(fp32 path).