Downloads · 30 days
11
15% of all-time downloads
procmarco/pina-500m
pina-500m is a machine learning model from procmarco. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A 536M-parameter decoder-only language model for Italian + English + code, trained from scratch. This is an early / partial base-model checkpoint — about 9.19B tokens (~23% of a planned 40B-token base run) — released…
Downloads · 30 days
11
15% of all-time downloads
All-time downloads
74
Public
Parameters
536M
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 100%
From the Hugging Face model README
A 536M-parameter decoder-only language model for Italian + English + code, trained from scratch. This is an early / partial base-model checkpoint — about 9.19B tokens (~23% of a planned 40B-token base run) — released for inspection.
This repo contains model weights only (model.safetensors, bf16) plus config and
inference code. It intentionally does not include optimizer state; resumable training
checkpoints live separately in procmarco/pina-500m-ckpt.
Current source checkpoint: step=17520, token_offset=9186050048.
Recent eval notes:
Custom SparseLM (not a 🤗 Transformers architecture): RMSNorm, RoPE, QK-norm,
z-loss, tied embeddings, ReLU-gated sparse FFN, and DeepSeek-style sparse attention
(DSA) in warm-up mode (dense flash + lightning-indexer distillation during training).
Inference here uses plain dense causal attention.
Config: d_model=1536, n_layers=16, ffn_hidden=4096, n_heads=24,
vocab_size=49152, max_seq_len=4096.
40% Italian (FineWeb-2 ita) / 40% English (FineWeb-Edu) / 20% code (github-code-clean).
Tokenizer: procmarco/ita-en-code-bpe-48k.
Documents were packed as <|bos|> ... <|eos|>.
Self-contained — the modeling code (modeling_pina.py) ships in this repo, so you only
need torch safetensors tiktoken huggingface_hub:
import json, pickle, torch
from safetensors.torch import load_model
from huggingface_hub import hf_hub_download
from modeling_pina import SparseLM, SparseLMConfig
REPO, TOK = "procmarco/pina-500m", "procmarco/ita-en-code-bpe-48k"
cfg = SparseLMConfig(**{k: v for k, v in json.load(open(hf_hub_download(REPO, "config.json"))).items() if k != "architecture"})
model = SparseLM(cfg).eval()
if torch.cuda.is_available():
model = model.cuda().to(torch.bfloat16)
load_model(model, hf_hub_download(REPO, "model.safetensors"))
enc = pickle.load(open(hf_hub_download(TOK, "tokenizer.pkl"), "rb"))
bos, eos = enc.encode_single_token("<|bos|>"), enc.encode_single_token("<|eos|>")
ids = [bos] + enc.encode_ordinary("L'Italia è un paese famoso per")
out = model.generate(ids, max_new_tokens=60, temperature=0.8, top_k=40, eos_id=eos)
print(enc.decode(out[1:]))
See example.py for a runnable version.