Downloads · 30 days
0
procmarco/ita-en-code-bpe-48k
ita-en-code-bpe-48k is a machine learning model from procmarco. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tiktoken. The card lists the license as mit.
A byte-level BPE tokenizer, vocab 49152 (48k), trained with rustbpe on a 40% Italian / 40% English / 20% code mix (~4B chars), so it is efficient across all three:
Downloads · 30 days
0
Access
Public
Updated Jun 30, 2026
Repo size
810 KB
Likes
0
Public
Click a slice to open those files.
.pkl1.2 MB · 75%
From the Hugging Face model README
A byte-level BPE tokenizer, vocab 49152 (48k), trained with
rustbpe on a 40% Italian / 40% English
/ 20% code mix (~4B chars), so it is efficient across all three:
ita_Latn (filtered)sample/100BT (int_score ≥ 3)Inference uses tiktoken — tokenizer.pkl
is a pickled tiktoken.Encoding. GPT-4 split pattern (cl100k). Italian stays
efficient (IT/EN share Latin script ⇒ shared merges) while English and code are
first-class.
33 special tokens, ids 49119–49151. Pretraining packs each doc <|bos|> … <|eos|>.
| group | tokens |
|---|---|
| core | <|bos|> (49119), <|eos|> (49120), <|pad|> (49121) |
| chat | <|system|>, <|user|>, <|assistant|>, <|observation|> |
| turn | <|eot|> |
| reasoning | <think>, </think> |
| tools | <tool_call>, </tool_call>, <tool_response>, </tool_response> |
| code FIM | <|fim_begin|>, <|fim_hole|>, <|fim_end|> |
| reserved | <|reserved_0|> … <|reserved_15|> |
tokenizer.pkl — pickled tiktoken.Encoding.token_bytes.pt — per-id UTF-8 byte length (0 for specials), for bits-per-byte eval.summary.json — training config + stats.import pickle
from huggingface_hub import hf_hub_download
enc = pickle.load(open(hf_hub_download("procmarco/ita-en-code-bpe-48k", "tokenizer.pkl"), "rb"))
ids = enc.encode_ordinary("def somma(a, b):\n return a + b # Ciao")
print(len(ids), enc.decode(ids))
pip install tiktoken runs it; training used rustbpe. Companion token dataset:
procmarco/ita-en-code-tokens-48k.