Downloads · 30 days
17
33% of all-time downloads
GPUburnout/GPUburnout-1B-90K
GPUburnout-1B-90K is a text generation model from GPUburnout. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A 1.04 billion parameter Llama-style language model trained from scratch on 11.8B tokens for $175.
Downloads · 30 days
17
33% of all-time downloads
All-time downloads
51
Public
Parameters
1B
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.1 GB · 100%
From the Hugging Face model README
A 1.04 billion parameter Llama-style language model trained from scratch on 11.8B tokens for $175.
| Benchmark | Metric | Score | Random |
|---|---|---|---|
| ARC-Easy | acc | 47.1% | 25% |
| HellaSwag | acc_norm | 28.8% | 25% |
| ARC-Challenge | acc_norm | 23.3% | 25% |
| MMLU | acc | 23.0% | 25% |
Includes ChatML special tokens for future SFT:
<|im_start|> (32000), <|im_end|> (32001)<|system|> (32002), <|user|> (32003), <|assistant|> (32004)from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("GPUburnout/GPUburnout-1B", torch_dtype="float16")
tokenizer = AutoTokenizer.from_pretrained("GPUburnout/GPUburnout-1B")
inputs = tokenizer("The capital of France is", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Full training journey documented at gpuburnout.com
Jun Park (@GPUburnout)