Downloads · 30 days
0
aethera-gp/kotodama-3b-base-final
kotodama-3b-base-final is a text generation model from aethera-gp. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A 3B parameter language model trained from scratch with Block Attention Residuals and NCA pre-pretraining.
Downloads · 30 days
0
Access
Public
Updated Jul 31, 2026
Repo size
9.4 GB
Likes
1
Public
Click a slice to open those files.
.zst9.4 GB · 100%
From the Hugging Face model README
A 3B parameter language model trained from scratch with Block Attention Residuals and NCA pre-pretraining.
This is the final base checkpoint: the full 384B-token schedule — 346B tokens at peak LR + 38B-token cosine cooldown to LR=0 (steps 175,780 → 195,311), completed 2026-06-09. It is the best checkpoint of the run on every internal and external evaluations.
Training compute provided by partnership with Anima Labs.
[0,1,3,7,15,19,24] — learned routing over depth at each sublayerThe exact training config is included in this repo as 3b-language.yaml.
lm-evaluation-harness 0.4.11, zero-shot:
| Task | chinchilla-66B | final-384B |
|---|---|---|
| HellaSwag (acc_norm) | 36.3 | 46.7 |
| PIQA (acc) | 64.5 | 68.4 |
| ARC-Easy (acc) | 51.6 | 55.1 |
| ARC-Challenge (acc_norm) | 24.4 | 27.3 |
| BoolQ (acc) | 58.4 | 61.9 |
| COPA (acc) | 68.0 | 71.0 |
| SciQ (acc) | 82.6 | 87.0 |
| Winogrande (acc) | 52.4 | 55.6 |
| LAMBADA (acc / ppl) | 38.2 / 23.4 | 49.7 / 11.1 |
| WikiText (word_ppl) | 26.08 | 17.75 |
UncheatableEval-2026-04 (bits-per-byte on post-cutoff data, 15 domains): mean 0.852 vs 0.980 (chinchilla) — wins all 15 domains. Strongest: arxiv/github (0.65–0.72); weakest: non-English (1.30–1.76). Cooldown isolation on near-token-matched checkpoints: the 6B cosine decay alone accounts for −5.7% mean BPB.
Evaluate in bf16. The training-time train/loss telemetry (fp8 + compile path) is a noisy estimator and not a reliable quality signal — it rose during the cooldown while the model improved on every held-out eval. All quality claims here are from bf16 evals.
This checkpoint requires the kotodama model code to load.
git clone https://github.com/LuxiaSL/kotodama.git
cd kotodama
# Serve interactively
python serve.py --checkpoint /path/to/step_00195311.pt.zst --model_size 3b --port 2222
# Then query:
curl http://localhost:2222/v1/completions \
-d '{"prompt": "The theory of everything", "max_tokens": 200, "temperature": 0.7}'
Or load the weights directly:
import io, torch, zstandard
raw = zstandard.ZstdDecompressor().stream_reader(open("step_00195311.pt.zst", "rb")).read()
ckpt = torch.load(io.BytesIO(raw), map_location="cpu", weights_only=False)
state_dict = ckpt["model"] # -> load into the model with the DD-3B config in 3b-language.yaml
Sampling note: use pure temperature sampling (no top-p) — top-p degraded quality in our evals.
Raw PyTorch checkpoint (.pt.zst, zstd-compressed, ~9.4GB; ~30GB decompressed). Contains model state dict, both optimizer states (Muon + AdamW), scheduler, and training metadata (including the fixed probe batch). The model code handles decompression automatically.
Apache 2.0