Downloads · 30 days
0
BEE-spoke-data/cl100k_base-mlm
cl100k_base-mlm is a machine learning model from BEE-spoke-data. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Dec 29, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json5.8 MB · 86%
From the Hugging Face model README
cl100k_base: as HF MLM tokenizerbased on RobertaTokenizerFast
from pathlib import Path
from transformers import RobertaTokenizerFast, AutoTokenizer
repo_id = "BEE-spoke-data/cl100k_base-mlm"
tk = AutoTokenizer.from_pretrained(repo_id)
len(tk)
# 100266
testing that it does what it should:
input_text = "i love memes"
tokenized_ids = tk.encode(input_text)
decoded_tokens = tk.convert_ids_to_tokens(tokenized_ids)
print(f"for input '{input_text}' -> {tokenized_ids} -> {decoded_tokens}")
# for input 'i love memes' -> [100277, 72, 3021, 62277, 100278] -> ['<s>', 'i', 'Ġlove', 'Ġmemes', '</s>']