Downloads · 30 days
16
6% of all-time downloads
iand666/makemore-mlp
makemore-mlp is a machine learning model from iand666. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A character-level MLP language model trained on the names dataset, built following Andrej Karpathy's makemore series.
Downloads · 30 days
16
6% of all-time downloads
All-time downloads
279
Public
Parameters
11.9K
47.9 KB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors47.9 KB · 89%
From the Hugging Face model README
A character-level MLP language model trained on the names dataset, built following Andrej Karpathy's makemore series.
The model learns to generate human-like first names by predicting the next character from the previous block_size characters.
C: maps each character to a learned emb_dim-dimensional vectorW1: linear + tanh, input size block_size * emb_dimW2: projects to vocab_size logits (27 classes: a–z + end token)| Hyperparameter | Value |
|---|---|
| block_size | 3 |
| emb_dim | 10 |
| hidden_dim | 200 |
| vocab_size | 27 |
Train loss: ~2.18 · Dev loss: ~2.20
import torch
import torch.nn.functional as F
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("iand666/makemore-mlp", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("iand666/makemore-mlp", trust_remote_code=True)
model.eval()
def gen_name(model, tok, block_size=3):
context = [0] * block_size
name = ""
with torch.no_grad():
while True:
x = torch.tensor([context])
logits = model(x)["logits"]
probs = F.softmax(logits, dim=-1)
nxt = int(torch.multinomial(probs[0], num_samples=1).item())
context = context[1:] + [nxt]
if nxt == 0:
break
name += tok.convert_ids_to_tokens(nxt)
return name
for _ in range(10):
print(gen_name(model, tok))