Downloads · 30 days
363
100% of all-time downloads
d0rj/q-51M-base
q-51M-base is a text generation model from d0rj. Use it when you need the model to write or continue text. It is set up for transformers.
A 50,878,208-parameter English base model in the Tiny llm ablation experiment. Trained from scratch on exactly 3,932,160,000 source tokens over 15,000 optimizer steps. The token count measures processed input blocks,…
Downloads · 30 days
363
100% of all-time downloads
All-time downloads
363
Public
Parameters
50.9M
204 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors204 MB · 97%
From the Hugging Face model README
A 50,878,208-parameter English base model in the Tiny llm ablation experiment. Trained from scratch on exactly 3,932,160,000 source tokens over 15,000 optimizer steps. The token count measures processed input blocks, not unique text or supervised target tokens.
10 decoder layers, width 512, SwiGLU width 1792, 8 query / 2 KV heads, RoPE, per-head QK normalization, gated residual branches, tied embeddings; context 2048.
Architecture and unchanged 32,768-token tokenizer: q-project/Q-50M-Base. This checkpoint was trained from random initialization; it is not a fine-tune of the reference weights.
sample-10BT, streamed from local Parquet shards; shuffle buffer 100,000.Full selected task splits, no added few-shot examples, lm-eval 0.4.12, no chat template, BF16 on RTX 5070 Ti, maximum context 2048 (ArithMark: 1024). Accuracy is a percentage. ± is one standard error; the separate bracketed column is the 95% Wilson confidence interval. Intervals describe finite evaluation-sample uncertainty, not variation across training seeds; no multiple-comparison correction is applied.
| Dataset | Split | Examples | Metric | Score ± SE (%) | 95% CI (%) |
|---|---|---|---|---|---|
| HellaSwag | validation | 10,042 | acc_norm | 29.18 ± 0.45 | [28.30, 30.07] |
| ARC-Easy | test | 2,376 | acc_norm | 43.31 ± 1.02 | [41.33, 45.31] |
| ARC-Challenge | test | 1,172 | acc_norm | 24.23 ± 1.25 | [21.87, 26.77] |
| PIQA | validation | 1,838 | acc_norm | 59.90 ± 1.14 | [57.64, 62.12] |
| WinoGrande | validation | 1,267 | acc | 50.04 ± 1.41 | [47.29, 52.79] |
| OpenBookQA | test | 500 | acc_norm | 28.20 ± 2.01 | [24.43, 32.30] |
| BoolQ | validation | 3,270 | acc | 59.88 ± 0.86 | [58.19, 61.55] |
| LAMBADA OpenAI | test | 5,153 | acc | 20.86 ± 0.57 | [19.77, 21.99] |
| ArithMark-3 | train | 1,000 | acc_norm | 36.70 ± 1.52 | [33.77, 39.73] |
| Balanced COPA | train | 1,000 | acc | 54.60 ± 1.58 | [51.50, 57.66] |
| CommonsenseQA | validation | 1,221 | acc | 19.66 ± 1.14 | [17.52, 21.98] |
| SciQ (with support) | test | 1,000 | acc_norm | 64.20 ± 1.52 | [61.18, 67.11] |
| TruthfulQA MC2 | validation | 817 | acc | 43.73 ± 1.51 | — |
| BananaMind Base 1.1 | test | 350 | raw_accuracy | 48.00 ± 2.67 | [42.82, 53.23] |
| MMLU continuation | test | 14,042 | acc | 25.44 ± 0.37 | — |
| BLiMP | train | 67,000 | acc | 77.27 ± 0.14 | — |
Q50M uses autoregressive continuation likelihood. LAMBADA accuracy requires the complete final-word token sequence. acc_norm is harness length-normalized option scoring; raw accuracy is also stored in results.json.
WikiText-2 raw test, conditional continuation: CPU FP32 re-evaluation on 291 nonoverlapping blocks (512 prefix + 512 scored suffix tokens), 148,992 scored tokens; 335 tail tokens excluded. NLL 3.560862, 95% CI [3.521807, 3.599507]; token PPL 35.194, 95% CI [33.846, 36.580]. Percentile block bootstrap, 10,000 resamples, seed 2026; exponentiate NLL endpoints for PPL. Blocks are the resampling unit; this does not model all within-document dependence. This is not standard rolling AR or word PPL. The earlier BF16 point is retained separately in TensorBoard, with no borrowed FP32 interval.
The metadata contains author-reported model-index scores. The evaluated dataset repositories had no registered eval.yaml on 2026-09-20, so no .eval_results leaderboard entry or verified badge is claimed. Machine-readable results and provenance.
Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled, no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.
Full results, provenance and group scores. Updated machine-readable results. TensorBoard events contain these new scores at step 15,000.
Install requirements.txt (tested with Transformers 5.17.0 / PyTorch 2.11.0). Custom model code is included; trust_remote_code=True is required. This example runs on CPU.
from transformers import AutoTokenizer, AutoModelForCausalLM
repo = "d0rj/q-51M-base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval()
inputs = tokenizer("The capital of France is", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
To reproduce the core evaluation from a downloaded repository, install evaluation/requirements.txt and run:
python evaluation/run_core.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output evaluation-rerun
To reproduce after downloading this model repository, accept the BananaMind dataset terms, authenticate with hf auth login, then run in a suitable CUDA environment:
pip install -r evaluation/comparison-20261001/repro/requirements.txt
python evaluation/comparison-20261001/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output comparison-rerun
The bundled runner uses the published model classes with the exact evaluation adapters and tokenizer. --limit produces smoke results only. Raw dataset examples are not included in this release.
TensorBoard event files include training telemetry and eval/<task>/<metric> at step 15,000, plus separate CI bounds. Recovered text-log telemetry covers steps 12,020–15,000 only (150 points); earlier training telemetry is unavailable. train/loss was corrected for the historical gradient-accumulation logging scale; recovery/loss_as_printed preserves the printed values.
These are small English continuation models, not instruction-tuned assistants. Equal source-token budgets do not imply equal target-token supervision or FLOPs. Benchmark contamination was not audited; results are from one training seed. Reference-model scores from different prompts, tokenizers or corpora are not directly interchangeable.