Downloads · 30 days
0
fs90/gemma-2-tokenizer-spm
gemma-2-tokenizer-spm is a machine learning model from fs90. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for splintr. The card lists the license as gemma.
Google's Gemma 2 tokenizer.model converted, id for id, into a plain-text format that a human can read, diff and audit — and that a tokenizer can load without a protobuf dependency.
Downloads · 30 days
0
Access
Public
Updated Aug 11, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.spm6.1 MB · 100%
From the Hugging Face model README
.spm)Google's Gemma 2 tokenizer.model converted, id for id, into a plain-text
format that a human can read, diff and audit — and that a tokenizer can load
without a protobuf dependency.
This is a modified file. See Modification notice
below and the NOTICE file. Use is subject to the Gemma Terms of Use,
including the Section 3.2 use restrictions — see Licence.
Nothing here is a model. This is a tokenizer vocabulary: the list of pieces a tokenizer splits text on, with the scores that decide merge order.
A SentencePiece tokenizer.model is a protocol buffer. That makes it opaque:
you cannot diff two of them, grep one, or see in a pull request what a change
did. It also means anything that wants to read one needs a protobuf parser and
the SentencePiece schema.
The obvious alternative — the .tiktoken format, base64(token) rank per line
— is lossy for SentencePiece, in three separate ways:
<0x41>; storing the raw byte forces a reader to reconstruct that
spelling by scanning for a run of 256 consecutive ids.USER_DEFINED pieces
verbatim, before merging; they are never merge candidates. CONTROL pieces
are never matched from text at all. Both score 0.0 and both are spelled
<...>, so neither the score nor the spelling tells them apart.That third one is not theoretical. Gemma 2 declares 245 USER_DEFINED pieces —
HTML markers such as <blockquote> and <table>. Drop the type and
<blockquote> is re-merged from < + blockquote + >, which measurably
mistokenized 5.6% of real documents in testing. Gemma 3, which declares
6,410 of them (it adds the whitespace and newline runs), was worse.
This format keeps all three.
One line per token id, in ascending id order, no gaps:
<base64 of the piece, UTF-8 encoded> <score> <type>
PHBhZD4= 0.0 3 # <pad> score 0.0 CONTROL
PGVvcz4= 0.0 3 # <eos> score 0.0 CONTROL
PGJvcz4= 0.0 3 # <bos> score 0.0 CONTROL
id_to_piece(i), so <0x41> keeps its real
byte-fallback spelling and ▁ word-boundary runs keep theirs. Base64 because
a piece may contain spaces, newlines or invalid-looking bytes.get_score(i), written as the shortest decimal that round-trips
the IEEE-754 value.ModelProto.SentencePiece.Type enum:
1 NORMAL, 2 UNKNOWN, 3 CONTROL, 4 USER_DEFINED, 6 BYTE.The id is the line's position, so ids cannot be duplicated or non-monotonic by construction — there is no id field to disagree with the ordering.
| pieces | 256,000 |
| NORMAL | 255,495 |
| USER_DEFINED | 245 |
| BYTE | 256 |
| CONTROL | 3 (<pad>, <eos>, <bos>) |
| UNKNOWN | 1 (<unk>) |
Two properties worth knowing before you write a loader:
add_dummy_prefix is false. Gemma does not prepend a word-boundary
marker to the input, unlike Llama and Mistral. Prepending one anyway shifts
the first piece of every input to a different token.byte_fallback is true, and the 256 <0xNN> pieces are how it is
reached.The conversion is checked in both directions before the file is written — every
piece, score and type is read back and compared against the source model, and
each score is round-tripped through f32 to confirm it survives a
single-precision parse. To repeat that yourself against your own copy of
Google's tokenizer.model:
python extract_spm_vocab.py --model tokenizer.model --output gemma2.spm --verify
The script is scripts/extract_spm_vocab.py in
splintr. Any SentencePiece implementation
will do the same job; the format is simple enough to re-derive in a few lines.
import base64
pieces, scores, types = [], [], []
for line in open("gemma2.spm"):
b64, score, kind = line.split()
pieces.append(base64.b64decode(b64).decode("utf-8"))
scores.append(float(score))
types.append(int(kind))
# USER_DEFINED (4) pieces are matched verbatim, never merged.
user_defined = {p for p, t in zip(pieces, types) if t == 4}
Extracted from Google's Gemma 2 tokenizer.model, MD5
f9e2445870ec741aa6346bbd75531bb4.
The vocabulary is Google's, not this repository's, and keeps Google's licence.
Required by Section 3.1 of the Gemma Terms of Use, and repeated in NOTICE:
gemma2.spm is a modified form of Google's Gemma 2 tokenizer.model — it
is not the original file. The protocol-buffer model was converted, id by id,
into the plain-text format described above. Nothing was added, removed,
reordered or rounded: all 256,000 pieces keep their ids, scores and piece
types, and no vocabulary entry differs from Google's file in any way.
Gemma is provided under and subject to the Gemma Terms of Use, found at
ai.google.dev/gemma/terms and reproduced in
full in the LICENSE file in this repository.
Use of this file is subject to the use restrictions in Section 3.2 of that
agreement. If you distribute this file, or anything derived from it, you must
pass those restrictions on to whoever you distribute to, provide them a copy of
the agreement, and include the NOTICE file.
Gemma 2, Gemma 3 and EmbeddingGemma are covered by these terms. Gemma 4 is not — it is released under Apache-2.0, under a separate licence at ai.google.dev/gemma/apache_2.