Downloads · 30 days
2.1K
53% of all-time downloads
episod/tt-tnt
tt-tnt is a text generation model from episod. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
tt-tnt is a 22M-parameter Llama-3-style base completion model with a 2048-token context. - It was trained from random initialization on one Tenstorrent Blackhole chip with ttml (tt-train), converted to Hugging Face fo…
Downloads · 30 days
2.1K
53% of all-time downloads
All-time downloads
3.9K
Public
Parameters
22M
911 MB on disk
Likes
0
Public
Click a slice to open those files.
.whl101 MB · 69%
From the Hugging Face model README
tt-tnt is a 22M-parameter Llama-3-style base completion model with a 2048-token context.
ttml
(tt-train), converted to Hugging Face format, and numerically verified against an independent
reimplementation.(1, 1)) through the Tenstorrent vLLM plugin, and it
also runs on CPU with plain transformers.Packaged and published with tt-model-manager 0.1.0 (manifest schema 6, v6 thin).
| Architecture | Llama-3 style: RoPE (θ=500000), RMSNorm, SwiGLU, grouped-query attention. 22,025,088 parameters, hidden 384, 6 layers, 6 heads / 3 KV heads |
| Hardware | One Blackhole chip, mesh (1, 1) (MESH_DEVICE=P150). Measured on one chip of a p300c in a TT-QuietBox 2. It hasn't been run on a physical P150 card |
| Context | 2048 tokens (max_model_len 2048). Longer prompts get HTTP 400. Ignore the model_max_length sentinel in tokenizer_config.json |
| Vocabulary | 32,000 (byte-level BPE, trained for this model) |
| Weights dtype | bfloat16 |
| License | Apache-2.0 (weights and code). The training corpus includes share-alike sources; see Licensing |
| Status | Functional: served on hardware, measured at concurrency 1 and 8 |
| Model CI v0 | not yet run |
Direct use:
Give it the opening of a simple story.
Out-of-scope use:
uv tool install tenstorrent # once — the Tenstorrent CLI, `tt`
tt model pull episod/tt-tnt
tt serve episod/tt-tnt
tt model pull installs the bundle into its own venv: pinned Python 3.12, ttnn==0.77.0,
the empty-target vLLM 0.25.1 build, and the TT vLLM plugin. It downloads the weights by
default, since tt has no --with-weights flag.
The server listens on port 20000 by default and walks upward if that port is busy. The first
serve converts the weights into a tensor cache under the install folder. On a TT-QuietBox 2, that
first serve reached a ready endpoint in 46 s, and a restart with the cache in place is faster.
The server is ready when the log prints Application startup complete.
Without tt-cli, tt-model alone does the whole job:
tt-model pull episod/tt-tnt --with-weights
tt-model serve episod/tt-tnt
Stop it with tt-model stop episod/tt-tnt. It opens one chip:
TT_VISIBLE_DEVICES, for example from a lease manager), run.sh uses the
first device of the grant.mesh-1x1.textproto lets the model open a single chip.| profile | hardware | mesh | max_num_seqs | block_size | max_model_len |
|---|---|---|---|---|---|
| P150 (default) | one Blackhole chip | (1, 1) | 32 | 64 | 2048 |
This is the only profile. (1, 1) is the only mesh shape this model can serve at: the
attention code requires both the head count (6) and the KV-head count (3) to divide the mesh
width, and no multi-chip preset has width 3.
It is a base completion model, so the completions endpoint is the natural one:
curl -s http://localhost:20000/v1/completions \
-H 'Content-Type: application/json' \
-d '{"model": "episod/tt-tnt",
"prompt": "Once upon a time, there was a little",
"max_tokens": 128, "temperature": 0.8, "top_p": 0.95}'
/v1/chat/completions works, but the model does not chat. Since 2026-09-27 the tokenizer
ships a plain chat template so that chat clients get a completion instead of an HTTP 400. The
template works like this:
A user message is therefore a story opening, and the reply is its continuation:
curl -s http://localhost:20000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "episod/tt-tnt",
"messages": [{"role": "user", "content": "Once upon a time, a little girl named Lily"}],
"max_tokens": 80, "temperature": 0}'
Served on 2026-09-27 (greedy, 80 tokens), that request returned: " was very hungry. She wanted to eat some food, but her mom said she had to eat her favorite food. Lily was very hungry and wanted to eat her favorite food. …" Every prompt token vLLM received matched a local render of the template, in 3 of 3 requests, one of them a three-message history.
Sampling at temperature 0.8 / top_p 0.95 is the representative way to read this model. Greedy decoding produces repetition loops (see Limitations).
On CPU, no Tenstorrent hardware is required:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("episod/tt-tnt")
model = AutoModelForCausalLM.from_pretrained("episod/tt-tnt").eval()
ids = tok("Once upon a time, there was a little", return_tensors="pt").input_ids
with torch.no_grad():
out = model.generate(ids, max_new_tokens=60, do_sample=True, temperature=0.8, top_p=0.95)
print(tok.decode(out[0], skip_special_tokens=True))
Greedy decoding, 60 new tokens, from the frozen evaluation set
(docs/measurements/samples-tt-tnt-v3.md):
Once upon a time, there was a little girl named Lily. She loved to play outside in the park. One day, she saw a big, shiny rock on the ground. She picked it up and showed it to her mom. "Look, Mommy! I found a shiny rock!" she said. Her mom smiled and said, "That
And one that stops on its own rather than being cut off at the limit:
The ants had learned that being eaten was a way of helping others. The moral of the story is that it's important to be kind to others and to help others.
Both come from near the top of the frozen set of 15 and aren't typical of it. The linked file
shows the full range, including the TinyStories collapses to "a little girl named Lily" and
several hard repetition loops. Sampled output at temperature 0.8 is in
docs/measurements/samples-tt-tnt-v3-t0.8.md
and gives the more representative read of the model's range.
Serving throughput and latency (on device, 1 chip, measured 2026-09-27):
| sampling | ISL / OSL | concurrency | N | TTFT median (p99) | TPOT median (p99) | tok/s per user | output tok/s |
|---|---|---|---|---|---|---|---|
| greedy | 128 / 128 | 1 | 16 | 5.11 ms (5.64) | 1.81 ms (1.97) | 551 | 539 |
| greedy | 128 / 128 | 8 | 64 | 14.53 ms (19.19) | 2.00 ms (2.08) | 501 | 3,836 |
| greedy | 1900 / 128 | 1 | 8 | 13.04 ms (14.38) | 2.95 ms (3.00) | 338 | 330 |
| default (temperature 1.0) | 128 / 128 | 1 | 16 | 6.48 ms (6.99) | 3.45 ms (3.81) | 290 | 286 |
| default (temperature 1.0) | 128 / 128 | 8 | 64 | 20.49 ms (26.06) | 6.42 ms (6.61) | 156 | 1,227 |
Methodology.
0000:01:00.0) of a p300c in a TT-QuietBox 2, mesh (1, 1).vllm bench serve over streaming /v1/completions, with random-token prompts
(seed 0, --ignore-eos) and 2 warmup requests.What the numbers mean.
temperature: 0
where greedy output is acceptable.0000:03:00.0) during tuning, against 1.77–1.81 ms on 0000:01:00.0. That wasn't
investigated.docs/measurements/serving-tune-2026-09-27.md).Accuracy (CPU, fp32, 2048-token window, on the same weights, sha256 97e19118…):
| metric | value | reference / chance | source |
|---|---|---|---|
HF conversion vs independent pure-NumPy ttml forward | max abs logit diff ~6e-6 (correlation ≈ 1 − 1e-13) | identical function | docs/model-development-troubleshooting.md, tests/test_to_hf.py |
StoryCloze (xstory_cloze en, 1,511 items) | 0.5705 | chance 0.5281 (measured class balance); context-blind 0.5222 | README.md StoryCloze section, docs/measurements/storycloze-tt-tnt.json |
| wikitext bits/byte | 1.4867 | GPT-2 small 0.9769 | lm-eval 0.4.9, docs/measurements/external-tt-tnt-v3.md |
| lambada_openai last-word acc | 0.0798 | GPT-2 small 0.3256 | same |
| piqa acc / arc_easy acc | 0.5326 / 0.2963 | chance 0.50 / 0.25 | same |
| hellaswag acc / arc_challenge acc / mmlu acc | 0.2584 / 0.1817 / 0.2292 | chance 0.25. At or below chance, so these aren't measurements of the model | same |
Terminates on </s> (frozen 15-prompt set, 128 max tokens) | 5/15 greedy, 11/30 sampled | previous checkpoint 0/15, 0/30 | docs/measurements/behaviour-tt-tnt-v1-vs-v3.md |
No on-device accuracy-vs-reference measurement (PCC or token agreement against CPU) has been recorded for tt-tnt.
Context is 2048 tokens, enforced.
max_model_len is 2048.max_num_seqs is 32, and that is a ceiling of the stack, not a tuning choice.
ValueError: Batch size 64 not supported). KV-cache capacity wasn't the
limit: 133,120 tokens are allocated.docs/measurements/serving-tune-2026-09-27.md.It can't serve on more than one chip. With 3 KV heads, only a (1, 1) mesh works, so it
can't use a P300 board's 2-chip or 4-chip meshes. It has been measured on one chip of a p300c,
not on a physical P150 card. There is no Model CI v0 run.
The tokenizer is pinned; the on-device weights load from main.
weights.revision to the commit that added the chat template, and vLLM
loads the tokenizer and template from that commit.episod/tt-tnt
without a revision, so it reads main. The two are identical today.The headline validation loss isn't comparable to the previous checkpoint's. 2.9937 against 4.2203 looks like a large gain, but most of it isn't one.
wikipedia_simple.docs/measurements/per-source-loss-tt-tnt-v1.md.
That number measured domain transfer, not learning.It has seen its training corpus once and hasn't memorized it.
artifacts/checkpoints-tt-tnt-v3/val_losses.jsonl) falls from 5.084 at
step 500 to ~3.28 by step 4,000.It can now stop, which no previous checkpoint could.
</s> tokens, while its
config.json declared eos_token_id: 2. Those checkpoints never ended generation on their own.</s>.TinyStories still dominates its voice. The frozen evaluation set
(docs/measurements/samples-tt-tnt-v3.md,
greedy, 15 prompts) shows the pattern:
Whether it uses all 2048 tokens is a separate question. The position-wise loss probe
(docs/measurements/context-use-tt-tnt-v3.md)
shows loss falling well past where the previous checkpoint went flat:
| positions | loss |
|---|---|
| [0,32) | 4.23 |
| [32,64) | 3.28 |
| [64,128) | 2.94 |
| [256,512) | 2.85 |
Past ~256 tokens, though, the improvement is about 0.02 nats per bucket against a standard error of ~0.08. That is directionally right and inside the noise.
Its RMSNorm layers did learn.
stochastic_rounding disabled. That silently
froze all 13 RMSNorm gammas at 1.0, bfloat16's rounding fixed point.stochastic_rounding: true.The model weights and this project's code are Apache-2.0.
The training corpus isn't. This checkpoint was trained on a nine-source blend. Two of those sources are share-alike:
tinystories
(roneneldan/TinyStories): 31% of
the blend, CDLA-Sharing-1.0.wikipedia_simple
(wikimedia/wikipedia): 15% of the
blend, CC-BY-SA-3.0.Full per-source licence, attribution, and the pinned dataset revisions are in
docs/corpus_licensing.md.
train/corpus.py), not written by
hand, so it can't drift from the registry.The corpus itself isn't redistributed. Only the recipe to reconstruct it byte-identically is
published, as episod/tt-tnt-corpus:
the source registry, pinned revisions, and fetch/prepare/measure/blend scripts.
Credit.
nanollama3 config and the ttml library.This model was originally published as tt-nanollama3.
ttml trainer, on TinyStories.train/configs/model/tt-tnt-384.yaml,
is a verbatim copy of tt-train's own nanollama3.yaml.tt-tnt.Three checkpoints have been published under this repo id:
| corpus | context | document separators | |
|---|---|---|---|
| the original (TinyStories-only) | TinyStories alone | 256 | none |
| the first blend-trained checkpoint | nine-source blend | 512 | none |
| this one | nine-source blend, separator-carrying revision | 2048 | yes |
| Corpus | Nine-source, licence-audited blend: TinyStories, Simple English Wikipedia, and seven curated Project Gutenberg slices (docs/corpus_blend.md). This is the first revision to carry </s> separators, averaging one per ~478 tokens |
| Tokens seen | 352,714,752: the full training split, one epoch |
| Steps | 10,764 at batch 16, sequence length 2048 (32,768 tokens/step) |
| Hardware | One Blackhole chip of a p300c (mesh_shape [1, 1]) in a TT-QuietBox 2. The other three chips were idle |
| Wall clock | ~91 minutes |
| Final train loss | 3.25 |
| Final validation loss | 2.9937 (end-of-run evaluate()). The periodic curve's last entry (step 10,764) reads 2.939; the two figures sample different held-out windows |
| Optimizer | AdamW, constant lr 3e-4, weight decay 0.01, stochastic_rounding: true |
This model grew out of the "Build an LLM from Scratch" lesson arc in tt-vscode-toolkit. It takes that arc past where the lessons stop: real training, checkpointing, conversion, and numerical verification.
docs/model-development-troubleshooting.md.| date | change |
|---|---|
| 2026-08-14 | First tt-kernel vLLM bundle pushed; current weights published (separator-carrying blend, 2048 context, HF commit e166888600) |
| 2026-08-15 | Bundle re-pushed; model-card frontmatter restored after a tagging bug |
| 2026-09-01 | Bundle manifest republished at schema v5 |
| 2026-09-05 | First v6 thin package |
| 2026-09-15 | Briefly a v5 fat (self-contained) package, then repackaged v6 thin (a69c5d8ba2); stale v5-fat tree removed |
| 2026-09-27 | Serving update; weights unchanged. The tokenizer gains a plain chat template, so /v1/chat/completions returns a continuation instead of HTTP 400. max_model_len is set to 2048, so over-length prompts get HTTP 400. The bundle ships mesh-1x1.textproto and uses the first chip of a grant, so it starts as one chip on a p300c. Adapter 1.1.0. The manifest records tt_metal_version 0.77.0 and pins the weights revision. Measured at concurrency 1 and 8, greedy and default sampling |
The two earlier checkpoints (TinyStories-only and first blend) and the tt-nanollama3 → tt-tnt rename predate this repo's recorded history, and their dates aren't recorded here.
episod/tt-tnt-1024 is the 123M sibling in the
same from-scratch family. It is a different model, not a hardware variant:
Question: … Answer: … format it was trained on;Neither package supersedes the other. Losses aren't comparable across the two (2048 vs 512 windows).
episod/tt-tnt-corpus is the corpus
recipe.
Questions or problems with this package: open a discussion at
https://huggingface.co/episod/tt-tnt/discussions. That is the one channel that reaches the
bundle's author.
A problem with the tt tooling itself: use tt report issue. It collects your environment
and opens a prefilled issue against tenstorrent/tt-cli, so it doesn't reach this package's
author.
Product feedback: send it to [email protected].
| component | built from |
|---|---|
| tt-metal | ttnn==0.77.0 (PyPI pin in requirements.txt; the manifest records tt_metal_version: 0.77.0) |
| tt-metal models code | tt-tnt-models-closure==0.77.0, a vendored models/common + models/tt_transformers/tt (not upstream tt-metal-models) |
| vLLM | 0.25.1, empty target (wheels/vllm-0.25.1+empty-…whl) |
| vllm-tt-plugin | 0.1.0 wheel built 2026-09-15 (sha256 3990ec4d…). Its source commit isn't recorded in the wheel |
| weights | model.safetensors sha256 97e191180d8e845c…, first published in HF commit e1668886007997ad0c2195b208e9c5da9c0a48ce, = artifacts/hf-tt-tnt-v3. The manifest pins weights.revision to the commit that added the chat template |
| adapter | tt_tnt_adapter.py 1.1.0 = bundle/tt_tnt_adapter.py at tsingletaryTT/tt-tnt commit 2830e43 |
| mesh | mesh-1x1.textproto = train/configs/mesh/mesh-1x1.textproto (tt-metal's p150 descriptor) |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager 0.1.0 (tt_kernel_version), integration build of the v6 thin fixes · schema 6 |