Downloads · 30 days
1.4K
41% of all-time downloads
SupraLabs/Supra2-Nano
Supra2-Nano is a text generation model from SupraLabs. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
1.4K
41% of all-time downloads
All-time downloads
3.3K
Public
Parameters
800K
3.2 MB on disk
Likes
23
Public
Click a slice to open those files.
.safetensors3.2 MB · 92%
From the Hugging Face model README

Supra2-Nano is an 800K-parameter language model trained from scratch on 1B tokens using the Qwen3 architecture. It is part of the Supra2 family, a set of small language models trained on consumer-grade hardware to study scaling behavior at extremely low parameter counts.
| Architecture | Qwen3 |
| Parameters | ~800K |
| Training tokens | 1B |
| Training steps | 7,000 |
| Precision | bfloat16 |
| Hardware | Kaggle 2x T4 |
| License | Apache 2.0 |
The model was trained with a data mix dominated by FineWeb-Edu, supplemented with a small proportion of Cosmopedia-v2 for synthetic instructional-style text.
Supra2-Nano uses the Qwen3 architecture at a heavily downscaled size: reduced hidden dimension, few attention heads, and a small number of transformer layers, all fit to an 800K parameter budget. Grouped-query attention and RMSNorm are used as in the original Qwen3 design. The tokenizer vocabulary is kept small (4,096 tokens) to avoid the embedding and output projection layers dominating the parameter count at this scale, a common failure mode when applying full-size tokenizers to tiny models.
At this parameter count, the model is not expected to produce coherent long-form generations. The purpose of this training run is to establish a baseline for the Supra2 family and to validate the training pipeline (tokenizer, architecture, data mix) before scaling to larger variants in the same family.
Zero-shot evaluation on standard small-LM benchmarks, compared against two other sub-1M/1M-parameter models from the same lab:

| Benchmark | Supra2-Nano (0.8M) | Supra-Mini-v6 (1M) | Supra-Mini-v3 (0.5M) |
|---|---|---|---|
| PIQA (acc_norm) | 0.53 | 0.54 | 0.50 |
| HellaSwag (acc_norm) | 0.27 | 0.27 | 0.25 |
| ARC-Easy (acc_norm) | 0.31 | 0.30 | 0.28 |
| ARC-Challenge (acc_norm) | 0.21 | 0.20 | 0.23 |
Scores across all three models sit close to random/majority-class baselines, which is expected at this parameter scale. PIQA is the strongest signal, consistent with what's typically observed in sub-million-parameter models — it has the lowest reasoning depth requirement of the four tasks. Supra2-Nano performs comparably to Supra-Mini-v6 despite having 20% fewer parameters, and outperforms Supra-Mini-v3 on three of four tasks.
import json
import os
import torch
from huggingface_hub import snapshot_download
from transformers import AutoModelForCausalLM, LlamaTokenizerFast
MODEL_NAME = "SupraLabs/Supra2-Nano"
PROMPT = "The history of artificial intelligence begins"
def load_tokenizer(model_name: str) -> LlamaTokenizerFast:
local_dir = snapshot_download(model_name)
config_path = os.path.join(local_dir, "tokenizer_config.json")
with open(config_path, "r", encoding="utf-8") as f:
cfg = json.load(f)
if isinstance(cfg.get("extra_special_tokens"), list):
cfg["extra_special_tokens"] = {}
with open(config_path, "w", encoding="utf-8") as f:
json.dump(cfg, f, indent=2, ensure_ascii=False)
return LlamaTokenizerFast.from_pretrained(local_dir)
def main():
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = load_tokenizer(MODEL_NAME)
model = AutoModelForCausalLM.from_pretrained(MODEL_NAME, torch_dtype=torch.bfloat16)
model.eval().to(device)
inputs = tokenizer(PROMPT, return_tensors="pt").to(device)
inputs.pop("token_type_ids", None)
with torch.no_grad():
output_ids = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.8,
top_p=0.9,
repetition_penalty=1.3,
pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
if __name__ == "__main__":
main()
This model is intended for research into small-model scaling, architecture ablations, and training pipeline validation. It is not intended for production use or any application requiring reliable text generation. Outputs at this scale will frequently be repetitive, ungrammatical, or incoherent.
If you use this model, please cite the SupraLabs organization on Hugging Face.