Downloads · 30 days
25
2% of all-time downloads
nanochat-students/base-d20
base-d20 is a machine learning model from nanochat-students. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model is trained with the nanochat recipe by Andrej Karpathy.
Downloads · 30 days
25
2% of all-time downloads
All-time downloads
1.2K
Public
Repo size
4.6 GB
Likes
3
Public
Click a slice to open those files.
.pt2.5 GB · 54%
From the Hugging Face model README
This model is trained with the nanochat recipe by Andrej Karpathy.
It was trained with a depth of 20 on 2 billion tokens and corresponds to this tokenizer. I will combine this repo with the tokenizer.
from transformers import AutoConfig, AutoModel, AutoTokenizer
import torch
model_dir = "nanochat-students/base-d20"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = AutoModel.from_pretrained(model_dir, trust_remote_code=True)
model = model.to(device)
model.eval()
tokenizer = AutoTokenizer.from_pretrained(model_dir, trust_remote_code=True)
prompt = "The capital of Belgium is "
input_ids = tokenizer.encode(prompt, prepend=tokenizer.get_bos_token_id())
ids = torch.tensor([input_ids], dtype=torch.long, device=device)
max_new_tokens = 50
with torch.inference_mode():
for _ in range(max_new_tokens):
outputs = model(input_ids=ids)
logits = outputs["logits"] if isinstance(outputs, dict) else outputs.logits
next_token = torch.argmax(logits[:, -1, :], dim=-1, keepdim=True)
ids = torch.cat([ids, next_token], dim=1)
decoded = tokenizer.decode(ids[0].tolist())
print(decoded)
timestamp: 2025-10-14 16:16:53
timestamp: 2025-10-14 16:11:41
timestamp: 2025-10-14 14:28:31
Logs are available on the trackio space here
<iframe src="https://nanochat-students-trackio.hf.space" frameborder="0" width="850" height="450" ></iframe>