Downloads · 30 days
0
wop/Cosmos-T2-Accelerate-Beta2
Cosmos-T2-Accelerate-Beta2 is a text generation model from wop. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<img src="https://calm-heart-d697.mmmmmm505090.workers.dev?text=Cosmos T2-Accelerate-beta" width="900" alt="Cosmos T2-Accelerate-beta" /
Downloads · 30 days
0
Access
Public
Updated Jun 3, 2026
Repo size
161 MB
Likes
0
Public
Click a slice to open those files.
.pt80.4 MB · 99%
From the Hugging Face model README
Universal Kaggle-ready training notebook for the Cosmos T2-Accelerate-beta series.
Notebook-generated card. Final metrics are filled after the Kaggle training run. This notebook is designed to stay Kaggle-friendly on 2x T4 GPUs. The goal is a reusable training recipe, not a production assistant.
| Model class | CosmosT2_Accelerate_LLM |
| Architecture | Decoder-only Transformer with RoPE, RMSNorm, SwiGLU, GQA, and a configurable Engram memory path |
| Parameters | ~9.96 M |
| Layers | 4 |
| Attention heads | 4 |
| KV heads | 1 |
| d_model | 64 |
| FFN hidden | 256 |
| Positional encoding | RoPE (rope_base=10000) |
| Normalization | RMSNorm |
| MLP | SwiGLU |
| Memory | Engram (use_engram=True, every 2 blocks) |
| Context length | 1028 |
| Training block size | 1028 |
| Tokenizer | Qwen/Qwen2.5-0.5B |
| Dataset | wop/XXXXXL-chain-of-thought |
| License | Apache-2.0 |
| Metric | Value |
|---|---|
| Rows used | 10,000 |
| Loss tokens seen | 10,591,118 |
| Epochs | 50 |
| Batch size | 6 |
| Peak LR | 3.00e-04 |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Wall-clock time | 26m 40s |
| Final training loss | 2.1199 |
| Final training perplexity | 8.33 |
| Final validation loss | 1.9287 |
| Final validation perplexity | 6.88 |
| Best validation loss | 1.8973 |
| Best epoch | 10 |
The notebook shows live loss and perplexity plots every 20 epochs and does not save the graph to disk.
import torch
from transformers import AutoTokenizer
from app import CosmosT2_Accelerate_LLM
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B")
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
ckpt = torch.load("$CHECKPOINT_NAME", map_location="cpu")
model = CosmosT2_Accelerate_LLM(**ckpt["config"])
model.load_state_dict(ckpt["model_state"])
model.eval()
prompt = tokenizer.apply_chat_template(
[
{"role": "system", "content": "Enable thinking features: INTUITION"},
{"role": "user", "content": "What is 12 * 7?"},
],
tokenize=False,
add_generation_prompt=True,
)
ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).input_ids
out = model.generate(ids, max_new_tokens=120, temperature=0.8, top_k=50)
print(tokenizer.decode(out[0], skip_special_tokens=False))
Use the Qwen2.5 chat template. The default system prompt is:
Enable thinking features: INTUITION
The model will then emit a <think> block followed by an answer when it has enough signal.
The model is trained to end its turn with the <|im_end|> token (ChatML), so generation stops there. During data prep, any example longer than the 1028-token context has its <think> reasoning replaced by a short placeholder (or is dropped) so every training sequence ends cleanly - the model is never trained on a mid-thought truncation.
This notebook is designed to train future Cosmos T2-Accelerate-beta variants by changing only the config block at the top.
@misc{cosmos-t2,
author = {wop},
title = {Cosmos-T2: A small from-scratch chain-of-thought Transformer},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/wop/Cosmos-T2-Accelerate-beta}
}