Downloads · 30 days
12
24% of all-time downloads
11-47/Nano.Deep.Reasoner.11m
Nano.Deep.Reasoner.11m is a machine learning model from 11-47. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
A small decoder-only causal language model trained from scratch.
Downloads · 30 days
12
24% of all-time downloads
All-time downloads
51
Public
Parameters
12.7M
568 MB on disk
Likes
0
Public
Click a slice to open those files.
.pt135 MB · 65%
From the Hugging Face model README
A small decoder-only causal language model trained from scratch.
<|input|><|think|><|thought|><|reasoning|><|answer|>True causal next-token prediction.
For an input sequence:
x[0], x[1], x[2], ...
the model learns:
x[0] -> x[1]
x[1] -> x[2]
x[2] -> x[3]
and so on.
Every individual training example is strictly limited to:
1095 content tokens + EOS
and padded to exactly 1096 positions.
No oversized example is intentionally split across separate training examples.
Starting from Session 3, new examples are selected via a deterministic shuffled scan of the dataset (seeded, reproducible across runs) rather than raw sequential order, to avoid category/source concentration within a session. Sessions 1-2 (40,000 examples) were selected sequentially before this correction and remain part of the trained corpus.
Plans11/Organized_PreTrain_1k_Context
Categories exposed by the dataset include:
Training state is persisted to Hugging Face, including:
Examples are tracked using SHA-256 content hashes.
The tokenizer becomes immutable after its initial creation.
This is an experimental small language model and is not guaranteed to produce factually or logically correct outputs.