Downloads · 30 days
152
25% of all-time downloads
prathamkode/particle-1.0
particle-1.0 is a text generation model from prathamkode. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
~100M-parameter Llama-style chat model trained from scratch (random init). Not a fine-tune of Llama, SmolLM, or any Hub base.
Downloads · 30 days
152
25% of all-time downloads
All-time downloads
613
Public
Parameters
110M
219 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors219 MB · 99%
From the Hugging Face model README
~100M-parameter Llama-style chat model trained from scratch (random init). Not a fine-tune of Llama, SmolLM, or any Hub base.
Weights are MIT. Training data still needs attribution (below).
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "prathamkode/particle-1.0"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
messages = [{"role": "user", "content": "hello"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, temperature=0.7)
print(tok.decode(out[0], skip_special_tokens=False))
Chat format:
<|user|>
hello
<|assistant|>
| Architecture | Llama-style decoder (RoPE, SwiGLU, RMSNorm, tied embeddings) |
| Parameters | ~100M (12 layers, 768 hidden, 12 heads) |
| Context | 2048 tokens |
| Tokenizer | Custom 32k byte-level BPE (not Llama / GPT-2 vocab) |
| Init | Random N(0, 0.02) — trained from scratch |
| Precision | BF16 training; Hub weights bfloat16 |
HuggingFaceFW/fineweb_edu_100BT-shuffled, first ~2B tokens.HuggingFaceTB/smol-smoltalk (first user/assistant turn + a few greeting seeds).SFT used that dataset as text only. No teacher model weights were copied.
Research / demo small chat model. Expect short replies, mistakes, and weak reasoning.