Downloads · 30 days
40
27% of all-time downloads
melephant/1-layer-addition
1-layer-addition is a text generation model from melephant. Use it when you need the model to write or continue text. It is set up for transformers.
Run s85nnxtf is a 1-block, bias-free causal transformer trained for 4-digit base-10 addition. Operands are zero-padded and answers use 5 digits, retaining overflow.
Downloads · 30 days
40
27% of all-time downloads
All-time downloads
148
Public
Parameters
43.6K
8 MB on disk
Likes
0
Public
Click a slice to open those files.
.pt7.6 MB · 91%
From the Hugging Face model README
Run s85nnxtf is a 1-block, bias-free causal transformer trained for
4-digit base-10 addition. Operands are zero-padded and answers use
5 digits, retaining overflow.
| Metric | Value |
|---|---|
| Validation loss | 0.003520 |
| Validation generated-token accuracy | 99.85% |
| Validation exact-answer accuracy | 99.32% |
| No-carry exact-answer accuracy | 97.27% |
| Single-carry exact-answer accuracy | 100.00% |
| Multiple-carry exact-answer accuracy | 98.44% |
| Carry-chain exact-answer accuracy | 97.27% |
unavailableThe complete resolved configuration, environment, metrics, source snapshot, and checkpoints are
available in training/. Machine-readable hashes and metrics are in
export_manifest.json.
This repository contains custom Transformers code. For reproducible or security-sensitive use, pin the commit revision printed by the uploader.
from transformers import AutoModelForCausalLM, AutoTokenizer
revision = "PINNED_COMMIT_HASH"
tokenizer = AutoTokenizer.from_pretrained(
"OWNER/REPO", trust_remote_code=True, revision=revision
)
model = AutoModelForCausalLM.from_pretrained(
"OWNER/REPO", trust_remote_code=True, revision=revision
)
inputs = tokenizer("0000 + 0000 =", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=model.config.answer_digits, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
This model is intended for mechanistic-interpretability research on its configured fixed-width addition task. It is not a general arithmetic system: inputs outside the configured grammar or width are unsupported, and generated answers must not be treated as reliable calculations.