Downloads · 30 days
140
13% of all-time downloads
CodeSoft/MetaDiffusion-150M-ChatBase
MetaDiffusion-150M-ChatBase is a text generation model from CodeSoft. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
MetaDiffusion-150M-ChatBase is a masked-diffusion language model converted from an autoregressive base model and chat-tuned for downstream experimentation. It uses bidirectional attention with timestep conditioning an…
Downloads · 30 days
140
13% of all-time downloads
All-time downloads
1.1K
Public
Parameters
170M
339 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors339 MB · 99%
From the Hugging Face model README
MetaDiffusion-150M-ChatBase is a masked-diffusion language model converted from an autoregressive base model and chat-tuned for downstream experimentation. It uses bidirectional attention with timestep conditioning and generates text through iterative masked denoising with left-to-right block commitment. Built by chat tuning CodeSoft/MetaDiffusion-150M-exp on smol-smoltalk (460K conversations, Apache-2.0) + no_robots (9.5K, Apache-2.0).
Downstream fine-tuning for specific tasks. This is NOT designed for production, instruction following is weak and coherent output is limited to roughly 60-100 tokens.
pip install -r scripts/requirements.txt
# Interactive, with live denoising view (--watch):
python scripts/chat.py --model-path . --watch
# One-shot:
python scripts/chat.py --model-path . --prompt "What is the capital of France?"
The --watch flag shows the response denoising in real time: step count,
noise level t, masks remaining, and the partial text building into place.
Generation defaults live in generation_config.json (128 steps, 96-token
block, temperature 0.7, repetition penalty 1.5).
# 1. Your data:
# a) HF dataset names (no_robots, alpaca, dolly, smol-smoltalk, math)
python scripts/prepare_data.py --model-path . \
--data-dir my_data --datasets no_robots
# b) Local ChatML files: .jsonl, .json, .parquet (messages/instruction shapes)
python scripts/convert_data.py --model-path . --input my_chat.jsonl \
--output my_data
python scripts/convert_data.py --model-path . --input ./data_folder \
--output my_data
# 2. Train
python scripts/train_chat.py --model-path . \
--data-dir my_data --output-dir my_checkpoints \
--lr 7e-5 --epochs 3 --patience 6
# 3. Chat with your model
python scripts/chat.py --model-path my_checkpoints/best.pt \
--tokenizer my_data/tokenizer --watch
# 4. Export a new release artifact (self-packages the scripts too)
python scripts/export_hf.py --checkpoint my_checkpoints/best.pt \
--tokenizer my_data/tokenizer --output ./MetaDiffusion-150M-MyTask --fp16
The whole pipeline runs on a single consumer GPU.
The model returns coherent sentences but loses coherence and factuality over multi-turn conversations.
<|im_end|> win at position 0 and produced empty responses on this
architecture at 150M; left-to-right fixed it (verified).<|im_end|> is committed.MetaDiffusion-150M-ChatBase is derived through the following process:
SupraLabs/Supra-1.5-50M-Base-exp.HuggingFaceTB/smol-smoltalk and HuggingFaceH4/no_robots.The resulting model uses bidirectional attention and timestep conditioning rather than conventional causal attention.
Apache-2.0