Downloads · 30 days
1.7K
59% of all-time downloads
episod/tt-tnt-1024
tt-tnt-1024 is a text generation model from episod. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
tt-tnt-1024 is a 123M-parameter Llama-3-style model with a 512-token context. It was trained from random initialization on Tenstorrent Blackhole with ttml (tt-train), in three passes:
Downloads · 30 days
1.7K
59% of all-time downloads
All-time downloads
2.8K
Public
Parameters
123M
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors246 MB · 70%
From the Hugging Face model README
tt-tnt-1024 is a 123M-parameter Llama-3-style model with a 512-token context. It was trained
from random initialization on Tenstorrent Blackhole with ttml (tt-train), in three passes:
databricks-dolly-15k question-answer
slice;It is served through the Tenstorrent vLLM plugin across all four Blackhole chips of a P300x2,
as a (1, 4) ring mesh with FABRIC_2D_TORUS_XY. Its chat endpoint uses the plain
Question: … Answer: … format it saw in training. It is small on purpose, and useful as an
instrument rather than a product. It is more fluent than earlier checkpoints, but not more
knowledgeable.
Status: Experimental. An earlier 4-chip output-quality regression didn't reproduce on these weights, but its cause was never identified (see Limitations).
It is the larger sibling of episod/tt-tnt, a project
first published as tt-nanollama3.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 6, v6 thin).
| Architecture | Llama-3 style decoder: 122,962,944 parameters, hidden 1024, 8 layers, 16 heads / 4 KV heads, SwiGLU (intermediate 2816), RoPE θ=500000, tied embeddings |
| Hardware | P300x2: 4 Blackhole chips (two p300c cards), (1, 4) ring, FABRIC_2D_TORUS_XY |
| Context | 512 tokens (max_model_len 512). Longer prompts get HTTP 400. The chat template keeps only the last 5 messages |
| Vocabulary | 32,000 (BPE, trained on this project's corpus) |
| License | Apache-2.0 (weights and code). The training data includes share-alike and attribution-licensed sources; see Licensing |
| Status | Experimental: serves on 4 chips and is measured at concurrency 1 and 8. The historical 4-chip regression wasn't reproduced, and its cause is unexplained |
| Model CI v0 | not yet run |
Direct use: a hardware-and-tooling research instrument. Use it for short story continuation,
short single questions in the Question: … Answer: … format, and experiments on training,
packaging, and serving on Tenstorrent hardware.
Out-of-scope use:
uv tool install tenstorrent # once — the Tenstorrent CLI, `tt`
tt model pull episod/tt-tnt-1024
tt serve episod/tt-tnt-1024
tt model pull installs the bundle into its own venv: Python 3.12, ttnn==0.77.0,
empty-target vLLM 0.25.1, and the TT vLLM plugin. It downloads the weights by default (tt
has no --with-weights flag). The server listens on port 20000 by default and walks upward if
that port is busy.
The first serve does the following:
Fabric initialized on 4 devices in the log);On a TT-QuietBox 2 the first serve reached a ready endpoint in 64 s, and a restart is faster.
The server is ready when the log prints Application startup complete.
Without tt-cli, tt-model alone does the whole job:
tt-model pull episod/tt-tnt-1024 --with-weights
tt-model serve episod/tt-tnt-1024
Stop it with tt-model stop episod/tt-tnt-1024. The bundle needs four chips:
TT_METAL_VISIBLE_DEVICES, that choice wins.TT_VISIBLE_DEVICES, run.sh uses the grant's first 4
devices. It refuses to start if the grant has fewer than 4.0,1,2,3.| profile | hardware | mesh | max_num_seqs | block_size | max_model_len |
|---|---|---|---|---|---|
| P300x2 (only profile) | 4 Blackhole chips, two p300c | (1, 4) ring (mesh-1x4-ring.textproto), FABRIC_2D_TORUS_XY | 32 | 64 | 512 |
4 KV heads divide 1, 2, and 4, so the architecture can shard across any of those. Only the 4-chip profile is packaged. On 2026-09-27, one chip (TP1) decoded at the same concurrency-1 speed as four (2.93 against 2.88 ms/token) during tuning. A 1-chip profile isn't packaged.
Use a vLLM TT plugin at or after c127c17. Earlier builds have a decode defect that degrades
free-running generation into repetition within a few tokens. The plugin reports version 0.1.0
either way, so a version check can't detect this; the bundle's adapter warns structurally.
The tokenizer's chat template (dolly_qa, since 2026-09-27) renders a conversation as the exact
token sequence of a databricks-dolly-15k document in the pretraining corpus:
Question: q Answer: a</s>.Question: q Answer:, and the model writes the answer.context field sat.curl -s http://localhost:20000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "episod/tt-tnt-1024",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
"max_tokens": 48, "temperature": 0}'
On 2026-09-27, this bundle on 4 chips returned the following (greedy; each reply ended on its
own with finish_reason: stop):
What is the capital of France? → The capital of France is Paris.
Who wrote Romeo and Juliet? → Juliet is the daughter of King Henry VIII of England.
(after the France exchange as history) What is the capital of Italy? → The capital of Italy is Rome.
The second answer is fluent and wrong, which is typical of this model. Every prompt token vLLM received matched a local render of the template, in 3 of 3 requests (one with a system message and a history).
Why this format. The previous template rendered Q: …\nAnswer:. That sequence occurs zero
times in the training tokens: the pipeline encoded the corpus line by line, so newlines were
never tokens. The results of the switch:
</s> 77% of the time, against 7% with
the old template.The proof is in
docs/measurements/chat-template-proof.json
and scripts/verify_chat_templates.py.
Why the 5-message cap exists. A growing conversation otherwise crashes the vLLM engine
outright. That is a generic tt-metal/vLLM defect, not this model: it reproduces identically on
stock meta-llama/Llama-3.2-1B-Instruct at the same context size (entry 6 in
docs/upstream-tt-metal-asks.md).
/v1/completions works for plain story continuation.
Serving throughput and latency (on device, 4 chips, measured 2026-09-27):
| sampling | ISL / OSL | concurrency | N | TTFT median (p99) | TPOT median (p99) | tok/s per user | output tok/s |
|---|---|---|---|---|---|---|---|
| greedy | 128 / 128 | 1 | 16 | 6.52 ms (7.45) | 3.16 ms (4.05) | 317 | 310 |
| greedy | 128 / 128 | 8 | 64 | 22.10 ms (105.50) | 3.40 ms (3.91) | 294 | 2,174 |
| greedy | 384 / 128 | 1 | 8 | 14.00 ms (20.78) | 2.96 ms (3.07) | 338 | 328 |
| default (temperature 1.0) | 128 / 128 | 1 | 16 | 7.46 ms (7.92) | 4.63 ms (5.71) | 216 | 212 |
| default (temperature 1.0) | 128 / 128 | 8 | 64 | 24.78 ms (46.45) | 7.88 ms (8.63) | 127 | 995 |
Methodology.
(1, 4) mesh,
FABRIC_2D_TORUS_XY.vllm bench serve over streaming /v1/completions, with random-token prompts
(seed 0, --ignore-eos) and 2 warmup requests.How to read these numbers:
docs/measurements/serving-tune-2026-09-27.md).Accuracy (CPU, on these exact weights, sha256 b2958652e9cc3308…):
| benchmark | this checkpoint (Stage B) | prior checkpoint (dialogue) | chance | source |
|---|---|---|---|---|
| wikitext bits/byte (lower is better) | 1.2551 | 1.4584 | n/a | docs/measurements/external-tt-tnt-1024-stageb.md |
| lambada_openai last-word acc | 0.2135 | 0.0980 | ~0 | same |
| piqa acc | 0.5925 | 0.5484 | 0.50 | same |
| arc_easy acc | 0.4272 | 0.3106 | 0.25 | same |
| hellaswag acc | 0.2745 | not recorded | 0.25 | same |
| winogrande acc | 0.4988 (at chance) | not recorded | 0.50 | same |
| arc_challenge acc | 0.2133 (below chance) | 0.1783 (below chance) | 0.25 | same |
| mmlu acc | 0.2292 (below chance) | 0.2295 | 0.25 | same |
| StoryCloze (1,511 items) | 0.6161 (context-blind 0.5248) | 0.6062 | 0.5281 | docs/measurements/storycloze-tt-tnt-1024-stageb.json, storycloze-tt-tnt-1024-vs-stageb.json |
Accuracy sources and methodology.
scripts/benchmark_external.py. StoryCloze comes from
scripts/eval_storycloze.py (xstory_cloze en, mean-per-token normalization).docs/current_model.json.How to read the accuracy rows:
episod/tt-tnt's 2048-window figures.On-device agreement with CPU (greedy, 10 prompts × 64 tokens, 2026-09-27; tokens that match CPU fp32 before the first divergence):
| topology | mean matching tokens | prompts identical to CPU |
|---|---|---|
4 chips (TP4, FABRIC_2D_TORUS_XY) | 17.4 | 4/10 |
| 1 chip (TP1) | 15.1 | 3/10 |
TP4 and TP1 are token-identical to each other on 4 of 10 prompts. Every divergence is ordinary bf16 drift into different, fluent English.
Context is 512 tokens, enforced.
max_model_len is 512.max_num_seqs is 32, and that is a ceiling of the stack, not a tuning choice.
tt_transformers 0.77 supports batch sizes 1, 2, 4, 8, 16, and 32 only. With 64 or 128 the
server fails to start (ValueError: Batch size 64 not supported). KV-cache capacity was not
the limit.
The only way past 32 is vLLM data parallelism. Four 1-chip replicas (DP4, 4 × 32 sequences) were measured:
| concurrency | DP4 | TP4 (shipped) |
|---|---|---|
| 1 (TPOT) | 3.25 ms (13% slower) | 2.88 ms |
| 8 | 1,018 tok/s | 2,487 tok/s |
| 32 | 3,713 tok/s | 7,779 tok/s |
| 128 | 12,366 tok/s | not measured |
DP4 only pays off at concurrency 128 and above, so it isn't shipped
(docs/measurements/serving-tune-2026-09-27.md).
4-chip output quality: a documented regression that did not reproduce.
FABRIC_2D_TORUS_XY), such as "Tryburg", "Alexandary", and
"Higheriq". The same prompt and sampling settings on 2 chips gave rough but recognizable
English
(docs/serving-with-tt-kernel.md
§8).
docs/upstream-tt-metal-asks.md).scripts/story_tools.py; now, greedy decoding).The tokenizer is pinned; the on-device weights load from main.
weights.revision to the commit that added the dolly_qa chat template,
and vLLM loads the tokenizer and template from that commit.episod/tt-tnt-1024
without a revision, so it reads main. The two are identical today.main without a bundle
rebuild. Any future weight upload would again change what installs serve, and it will come
with a repackage and a changelog entry.It repeats. Greedy decoding often falls into a repetition loop after the first sentence,
though the new template lets short answers stop on </s>.
tt-tnt-1024a's, at 3.32× the seed floor:| signal | delta | vs seed floor | verdict |
|---|---|---|---|
| 4-gram repeat rate | +0.0074 | 3.32× | worse |
| termination rate | −0.0076 | 0.52× | not interpretable |
| genre collapse | −0.0035 | 0.06× | not interpretable |
| loss at matched window | +0.0102 | — | no floor for this instrument |
docs/measurements/evaluation-tt-tnt-1024a-vs-tt-tnt-1024-dialogue.md).It isn't instruction-tuned beyond a 2% slice of databricks-dolly-15k.
Tool calling isn't supported on these weights. They emit no tool calls. A separate, unpublished continued-training checkpoint does (see History).
databricks-dolly-15k dialogue slice (2%, rendered as plain Question: … Answer: …
prose with no role markers) is CC-BY-SA-3.0.docs/corpus_licensing.md.HuggingFaceFW/fineweb-edu (sample-10BT, ODC-By 1.0,
pinned revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9).
episod/tt-tnt-corpus recipe.docs/current_model.json's corpus.note.| stage | data | steps | notes |
|---|---|---|---|
| dialogue (base) | curated ten-source blend (tokens-v4, 399,486,992 emitted tokens) | 10,764 | batch 64, seq 512 |
| Stage A | FineWeb-Edu, 2,529,270,500 tokens (the Chinchilla-matched budget for 123M) | 76,503 | cosine lr, --ddp 4 |
| Stage B | curated blend again (tokens-v4), one epoch | 10,761 | cosine lr warm restart, stochastic_rounding: True, final step 87,264 |
Training ran as 4-chip DDP across both p300c cards of a TT-QuietBox 2. The final validation loss
is 2.5373 (Stage B, real held-out loss over the full validation split). The ten-source blend is
episod/tt-tnt's nine sources plus the dolly slice.
⚠️ Everything in this section was measured against the prior (2026-08-20 to 2026-08-29 dialogue) checkpoint, not against the currently published Stage A/B weights. It records this model's lineage and makes no claims about the current weights.
Routing by physical die address is nearly free. Tokens can be routed to experts by where they live on the harvested 11×10 Tensix grid rather than by a learned gate:
Sparse routing (Mixture of Enthusiasts) beats dense from scratch. Both arms trained one epoch from init, paired on seed 5489:
Tool calling works structurally, and only structurally. A continued-training run teaches
four tools (factual_response, witty_response, absurdist_response,
misunderstood_question). They are emitted as <tool_call> blocks, which vLLM's hermes parser
turns into structured tool_calls:
| gate | trained | control |
|---|---|---|
| emits a tool call | 100% | 0% |
| parses | 85.9% | 0% |
| schema-valid | 75.0% | 0% |
| distinct tools used | 4/4 | 0/4 |
tool_calls end-to-end through the server.docs/measurements/tool-calling-stage{1,2}.json.
Its Q: …\nAnswer: format was the chat template this repo shipped until 2026-09-27.A five-slot think-block can be learned, but it doesn't help yet. The fine-tuned model emits
offer / accept / add / stakes / handback blocks:
stochastic_rounding defaults off on the SFT path). With the gammas free, 0% became 98%.episod-log.md,
2026-08-21.The reach dial (2026-08-23/24) works, is small, and is paused. Forcing a reach slot to
near / mid / far moves the realised semantic distance of the add word monotonically,
scene-paired over 826 held-out scenes:
| contrast | raw | frequency-residualised |
|---|---|---|
near < mid | +0.0839 (t 16.2) | +0.0324 (t 7.4) |
mid < far | +0.0456 (t 13.9) | +0.0281 (t 9.0) |
near < far | +0.1295 (t 23.3) | +0.0604 (t 12.5) |
reach slot shows nothing.add slot-hit shortfall 0.0896, a
near-side dip.docs/measurements/reach-dial.json through
scripts/eval_reach.py --rescore-from, with no model, tokenizer, or device.| date | change |
|---|---|
| 2026-08-18 | First published (dialogue-slice 512-context weights); first 1024-size tt-model bundle |
| 2026-08-29/30 | A 2048-context retrain was published (HF 038d6c6a8d), found worse at Q&A, and reverted within the hour to the 512 dialogue weights (57224f4b87). The chat template's 5-message guard was added |
| 2026-09-01 | Bundle manifest republished at schema v5 |
| 2026-09-05 | First v6 thin package |
| 2026-09-15 | Briefly a v5 fat package, then repackaged v6 thin (136d240427); stale v5-fat tree removed |
| 2026-09-25 | Weights replaced with the Stage A + Stage B checkpoint (9abed7c785), designated 2026-09-24. Bundle not rebuilt |
| 2026-09-27 | Repackaged for serving; weights unchanged. Changes: the chat template becomes dolly_qa (Question: … Answer: …</s>, the trained format; it replaces Q:\nAnswer:); max_model_len is set to 512 and adapter 1.1.0 clamps prefill buckets, so prompts of 129–512 tokens no longer crash the server and longer ones get HTTP 400; run.sh honours a chip grant; the manifest records tt_metal_version 0.77.0 and pins the weights revision. Hardware-verified on 4 chips; measured at concurrency 1 and 8, greedy and default sampling; 1-chip vs 4-chip output compared |
episod/tt-tnt is the 22M sibling in the same
from-scratch family. It is a different model, not a hardware variant: hidden 384, 3 KV
heads, 2048 context, no FineWeb-Edu, and no dialogue data, so its chat template is plain
continuation. It serves on one chip. Neither package supersedes the other.episod/tt-tnt-corpus is the curated
blend recipe. Stage A's FineWeb-Edu isn't part of it.https://huggingface.co/episod/tt-tnt-1024/discussions. That is the one channel that reaches
the bundle's author.tt tooling itself: use tt report issue. It opens a prefilled issue
against tenstorrent/tt-cli, not this package.The full build log and every measurement are at tsingletaryTT/tt-tnt.
| component | built from |
|---|---|
| tt-metal | ttnn==0.77.0 (PyPI pin in requirements.txt; the manifest records tt_metal_version: 0.77.0) |
| tt-metal models code | tt-tnt-models-closure==0.77.0, a vendored models/common + models/tt_transformers/tt (not upstream tt-metal-models) |
| vLLM | 0.25.1, empty target |
| vllm-tt-plugin | 0.1.0 wheel built 2026-09-15 (sha256 3990ec4d…). Its source commit isn't recorded in the wheel |
| mesh | mesh-1x4-ring.textproto (dims: [1, 4], [LINE, RING]), FABRIC_2D_TORUS_XY |
| weights | model.safetensors sha256 b2958652e9cc3308…, Stage B step 87,264, first published in HF commit 9abed7c7855d39583bb757a8bbf679ce569c40f6. The manifest pins weights.revision to the commit that added the dolly_qa chat template |
| adapter | tt_tnt_adapter.py 1.1.0 = bundle/tt_tnt_adapter.py at tsingletaryTT/tt-tnt commit 2830e43 |
| designation | docs/current_model.json, commit 44d8496 |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager 0.1.0, integration build of the v6 thin fixes · schema 6 |