Downloads · 30 days
3
16% of all-time downloads
Grayblock-AI/POET-124M
POET-124M is a machine learning model from Grayblock-AI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for sphere-attention. The card lists the license as mit.
124M-parameter POET (GPT-2 small size) — the scaling-trend checkpoint: the advantage grows, not shrinks, with size.
Downloads · 30 days
3
16% of all-time downloads
All-time downloads
19
Public
Repo size
495 MB
Likes
0
Public
Click a slice to open those files.
.pt495 MB · 100%
From the Hugging Face model README
124M-parameter POET (GPT-2 small size) — the scaling-trend checkpoint: the advantage grows, not shrinks, with size.
Second rung of the scaling study, parameter-matched to a GPT-2-style baseline (123.62M vs 124.01M total; 85.02M non-embedding each), same 983M tokens and schedule.
| scale | baseline best ppl | POET | paired margin |
|---|---|---|---|
| 51M | 2.938 | 2.828 | −2.9% (z = −68) |
| 124M (this model) | 2.72 | 2.75 final / 2.60 EMA | −3.3% (z = −74) |
The margin held and slightly grew across a 2.4x size jump — the opposite of the usual pattern for efficient-attention variants, which tend to fade with scale. A 354M rung is in progress.
Recipe: dim 768 / depth 12 / heads 12, seq 512, lr 6e-4, 60k steps, selective weight decay, EMA 0.999.
POET (POisson attEntion Transformer) replaces softmax attention with the closed-form spherical Poisson kernel. Queries and keys are L2-retracted onto the unit hypersphere and scored by
P_r(<q,k>) = (1 - r^2) / (1 - 2 r <q,k> + r^2)^(d/2)
with a learnable per-head radius r (resolution, replacing
temperature), rotary positions (rotations are isometries of the sphere), and
a QUEST-style query-norm sharpness term. Equivalently, each attention weight
is the harmonic measure of where a random walk started at the query's
interior point first exits the sphere at that key.
Because keys live on a sphere, attention admits a geodesic cap decomposition that supports budgeted sparse attention with closed-form per-query error certificates.
Architecture is otherwise a standard pre-LN transformer (GELU MLP, tied
embeddings, GPT-2 BPE). Custom code — not transformers-compatible.
pip install git+https://github.com/Grayblock-AI/spherical-attention
import torch, tiktoken
from sphere_attention.hub import load_poet
model = load_poet("Grayblock-AI/POET-124M")
enc = tiktoken.get_encoding("gpt2")
ids = torch.tensor([enc.encode_ordinary("def quicksort(arr):")])
out = model.generate(ids, max_new_tokens=64, top_k=20)
print(enc.decode(out[0].tolist()))
codeparrot/codeparrot-clean (permissively licensed Python), GPT-2 BPE.
The long-context variants use repository-grouped packing: files from the
same repo are packed contiguously so long sequences contain genuine
cross-file structure (imports, call sites, definitions).
POET-51M-certified for the regularized variant and the
gating mechanism that bounds the tail operationally.@misc{poet2026,
title = {POET: Poisson Attention Transformers with Certified Sparsity on the Hypersphere},
author = {Jerge, Michael},
year = {2026},
url = {https://github.com/Grayblock-AI/spherical-attention}
}
{
"model": "sphere-ff",
"dim": 768,
"depth": 12,
"heads": 12,
"seq_len": 512,
"batch_size": 32,
"steps": 60000,
"warmup": 1000,
"lr": 0.0006,
"solver_iters": 4,
"solver_tol": 0.001,
"eval_every": 1000,
"seed": 0,
"grad_accum": 2,
"no_bf16": false,
"r_init": 0.7,
"logit_scale": 10.0,
"tag": "125m",
"save_checkpoint": true,
"checkpoint_every": 0,
"s3_prefix": "",
"data_dir": "data1b",
"wandb": false,
"wandb_project": "spherical-attention",
"wandb_entity": null,
"cloudwatch": false,
"cloudwatch_region": "us-east-2",
"resume": "",
"ngpt": true,
"selective_wd": true,
"grad_steps": 1,
"ema": 0.999,
"cluster_reg": 0.0,
"abort_divergence": 1.5
}