Downloads · 30 days
23
21% of all-time downloads
Asilarkness/cubic-150m
cubic-150m is a text generation model from Asilarkness. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A 157M parameter bilingual research language model built on a custom attention block: a differentiable depth-memory path that lets every layer read a compressed copy of the state at every earlier layer, plus hierarchi…
Downloads · 30 days
23
21% of all-time downloads
All-time downloads
108
Public
Parameters
157M
942 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors314 MB · 99%
From the Hugging Face model README
A 157M parameter bilingual research language model built on a custom attention block: a differentiable depth-memory path that lets every layer read a compressed copy of the state at every earlier layer, plus hierarchical group retrieval in the two penultimate attention layers.
Training stage of this checkpoint: complete (dpo done) (step 0).
The table below is not an architecture comparison. This model saw 3 billion pretraining tokens. SmolLM2-135M saw roughly 2 trillion — a 600x larger data budget on a model of the same size. Pythia-160M saw ~300B. Any gap in the table is dominated by that difference, not by the attention design. Read it as "what 3B tokens buys you", nothing more.
All rows were produced by the same script, the same prompt formats and the same log-likelihood protocol in a single run, so the rows are comparable to each other. They are not directly comparable to published numbers from other harnesses, which differ in prompt formatting, few-shot count and normalization.
MMLU at this scale is at chance (25%) for every model in the table, this one included. It is reported because it was requested, not because it is informative below roughly 1B parameters.
| Model | Params | hellaswag | arc_easy | arc_challenge | piqa | winogrande | openbookqa | sciq | boolq | lambada_openai | mmlu | Avg | Δ chance |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CubicHierLM-157M (base format) | 157M | 30.1 | 39.1 | 25.3 | 57.8 | 50.8 | 25.2 | 70.4 | 61.8 | 20.1 | 27.1 | 40.8 | +10.8 |
| CubicHierLM-157M (chat, direct prompt) | 157M | 29.7 | 36.6 | 24.1 | 57.1 | 49.1 | 29.2 | 66.5 | 43.7 | 0.7 | 25.7 | 36.2 | +6.2 |
| CubicHierLM-157M (chat, reasoning prompt) | 157M | 30.1 | 35.7 | 24.0 | 57.1 | 50.8 | 28.8 | 65.9 | 42.2 | 1.1 | 26.3 | 36.2 | +6.2 |
| HuggingFaceTB/SmolLM2-135M | 135M | 43.9 | 59.9 | 29.2 | 68.2 | 53.3 | 32.4 | 79.9 | 61.1 | 41.1 | 23.1 | 49.2 | +19.2 |
| EleutherAI/pythia-160m | 162M | 29.0 | 37.7 | 24.7 | 58.7 | 50.7 | 25.2 | 63.6 | 43.9 | 10.9 | 24.1 | 36.9 | +6.9 |
| facebook/opt-125m | 125M | 30.3 | 40.9 | 22.7 | 61.8 | 50.4 | 27.4 | 68.9 | 55.8 | 38.1 | 24.2 | 42.1 | +12.1 |
| random chance | — | 25.0 | 25.0 | 25.0 | 50.0 | 50.0 | 25.0 | 25.0 | 50.0 | 0.0 | 25.0 | — | +0.0 |
Headline metric per task: acc_norm for multiple choice with unequal-length
options, acc otherwise. Evaluated on an evenly strided subsample of up to 1500 examples per task.
Read Δ chance, not Avg. The tasks have different random-guess floors —
0% for LAMBADA, 25% for the four-way questions, 50% for the binary ones — so a
plain average across them is not a meaningful quantity. Δ chance is the mean
margin over the random baseline and is the column that actually compares.
Two further caveats. BoolQ's validation set is about 62% "yes", so any score near 62 means the model is answering yes to everything and carries no signal. ARC-Challenge, OpenBookQA, WinoGrande and MMLU sit at chance for every model of this size and should not be read as differences.
This model is scored three times, because it went through instruction SFT and DPO and therefore has a chat format the baselines do not have:
Question: ... Answer: with no chat tokens. This is
the row comparable to the baseline models.<bos>, <|system|>,
<|user|>, <|assistant|>) with the concise-answer system prompt.<think> system
prompt.The reasoning row is not a chain-of-thought evaluation. Every task here is scored by the log-probability of each answer option; nothing is generated, so no reasoning trace ever exists. That row measures only whether the reasoning system prompt shifts which answer the model prefers. Measuring actual chain-of-thought would require generating traces, which this architecture has no KV cache for.
This is not a Transformers-compatible checkpoint. The architecture lives in
modeling_cubic.py in this repository.
python inference_example.py
Small, experimental, trained on a compute-optimal but tiny token budget. It will confabulate. Do not use it for anything factual, safety-relevant or mathematical without independent verification. Bilingual EN/RU by design, but Russian coverage is 13% of pretraining and correspondingly weaker.
Raw metrics: hellaswag=30.1, arc_easy=39.1, arc_challenge=25.3, piqa=57.8, winogrande=50.8, openbookqa=25.2, sciq=70.4, boolq=61.8, lambada_openai=20.1, mmlu=27.1