Downloads · 30 days
7K
40% of all-time downloads
YNSScarSaiyan/simi-weights
simi-weights is a text generation model from YNSScarSaiyan. Use it when you need the model to write or continue text. It is set up for flax. The card lists the license as apache-2.0.
Public weights-only mirror. No architecture source, no trainer, no private recovery tree.
Downloads · 30 days
7K
40% of all-time downloads
All-time downloads
17.7K
Public
Repo size
827 GB
Likes
0
Public
Click a slice to open those files.
Other827 GB · 100%
From the Hugging Face model README
Public weights-only mirror. No architecture source, no trainer, no private recovery tree.
SImi-2B is a custom Llama-style causal LM trained in JAX/Flax on an AMD MI300X. This repo is not a Transformers checkpoint. AutoModel.from_pretrained will not load it.
Created by Dakuwon Moody.
| Item | Value |
|---|---|
| Public role | Weights only. No source. |
| Student | SImiModel / SImi-2B |
| Parameters | ~2.32B (2,321,856,000 config estimate) |
| Live graph | nn.scan over 30 layers |
| Tokenizer | gpt2 (50,257) |
| Training mix | Full Hugging Face agent-grade streams (no toy slices) |
| Hardware | AMD Instinct MI300X, bf16, remat off, micro-batch 3, seq 1024 |
| Optimizer | Adafactor on the live run (step 99000 on disk is the older Adam snapshot) |
| Save cadence | every 5000 steps under distill/step_<n>/ |
| Latest public tree | distill/step_180000/ |
| Hosted inference | Not supported |
| Path | What |
|---|---|
distill/step_99000/ | Canonical unrolled restore. 30 layer_* blocks, Adam/MultiSteps Orbax tree. This is the known-good init. |
distill/step_100000/ .. later | Scanned Adafactor Orbax trees (scanned_layers). Written by the live MI300X loop. |
Local copies are deleted only after the Orbax blobs exist here.
Restore notes (this ROCm/Orbax stack):
layer_0..layer_29 into nn.scan.StandardRestore can fail (Layout / TensorStore). A metadata-shaped numpy restore of the params works. Fresh Adafactor opt is fine; numpy drops optax NamedTuples.metadata.json step number as proof a folder exists.Llama-style decoder-only:
| Property | Value |
|---|---|
| Vocabulary | 50,257 (GPT-2 BPE) |
| Hidden width | 2,560 |
| Depth | 30 layers |
| Query heads | 20 |
| KV heads | 4 |
| Head dimension | 128 |
| SwiGLU intermediate | 6,912 |
| Max sequence | 1,024 |
| RoPE theta | 10,000 |
| RMSNorm eps | 1e-5 |
| Compute / param dtype | bfloat16 |
| Estimated parameters | 2,321,856,000 |
| Property | Value |
|---|---|
| Tokenizer | gpt2 |
| Vocab size | 50,257 |
| BOS / EOS / PAD | 50256 |
| Packing | 1,024 tokens |
Same full HF streams as Uni and Vegeta: