Downloads · 30 days
165
19% of all-time downloads
dlab-spp/t0-mt-3b-instruct
t0-mt-3b-instruct is a text generation model from dlab-spp. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Type: instruction-tuned model (base model + persona-binding supervised fine-tuning).
Downloads · 30 days
165
19% of all-time downloads
All-time downloads
872
Public
Parameters
3B
29.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.9 GB · 100%
From the Hugging Face model README
Type: instruction-tuned model (base model + persona-binding supervised fine-tuning).
Trained with SPP from token zero plus reflection-focused midtraining, then post-trained with persona-binding SFT.
Synthetic Persona Pretraining (SPP) installs a target value persona during pretraining rather than only during alignment. Value-laden, first-person reflections, generated against a constitution, are appended to a subset of pretraining documents after a special <assistant> token. Attention masking and RoPE position aliasing keep the reflection from changing the continuation of the original document. This model is trained with SPP.
Base counterpart: dlab-spp/t0-mt-3b-base.
<assistant> marker token (vocabulary 49280).[N.M] citations; response-only loss, one epoch.There is no system prompt. Each assistant turn opens with <|im_start|><assistant>. Use the built-in chat template:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
repo = "dlab-spp/t0-mt-3b-instruct"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")
msgs = [{"role": "user", "content": "How should I think about honesty?"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=512)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=False))
This model is one point on a safety-data sweep. main is the default 10% mixture; the other fractions are published as revisions on this repo, so each can be loaded by passing revision=:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
repo = "dlab-spp/t0-mt-3b-instruct"
tok = AutoTokenizer.from_pretrained(repo) # identical at every revision
model = AutoModelForCausalLM.from_pretrained(
repo, revision="safety-60", dtype=torch.bfloat16, device_map="auto"
)
| Revision | Safety fraction | Safety examples | Instruct examples |
|---|---|---|---|
safety-0 | 0% | 0 | 300,000 |
safety-5 | 5% | 15,000 | 285,000 |
safety-10 — default, same weights as main | 10% | 30,000 | 270,000 |
safety-30 | 30% | 90,000 | 210,000 |
safety-60 | 60% | 180,000 | 120,000 |
Every mixture is 300,000 examples total, one epoch, response-only loss; safety prompts come from WildJailbreak and WildGuardMix and instructions from WildChat-1M. Only the ratio changes.
Research on alignment and safety (constitutional alignment, value generalization, jailbreak robustness). A research artifact, not a production model; it can produce incorrect or unsafe content.
@misc{minder2026syntheticpersonapretrainingalignment,
title={Synthetic Persona Pretraining: Alignment from Token Zero},
author={Julian Minder and Viktor Moskvoretskii and Raghav Singhal and Difan Jiao and Andy Arditi and Shaobo Cui and Yiderigun Borjigin and Kartik Bali and Stefan Krsteski and Harsh Raj and Huu Nguyen and Jannik Brinkmann and Ashton Anderson and Roland Aydin and Robert West},
year={2026},
eprint={2608.13482},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2608.13482},
}
License: to be finalised.