Downloads · 30 days
0
SceneWorks/qwen-image-tokenizer
qwen-image-tokenizer is a machine learning model from SceneWorks. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tokenizers. The card lists the license as apache-2.0.
A derived artifact for running Qwen/Qwen-Image on the native-Rust/MLX mlx-gen engine (SceneWorks).
Downloads · 30 days
0
Access
Public
Updated Jun 20, 2026
Repo size
11.4 MB
Likes
0
Public
Click a slice to open those files.
.json14.8 MB · 90%
From the Hugging Face model README
tokenizer.json)A derived artifact for running Qwen/Qwen-Image on the native-Rust/MLX mlx-gen engine (SceneWorks).
Qwen/Qwen-Image ships its Qwen2 BPE tokenizer as vocab.json + merges.txt only — there is
no fast tokenizer.json in the upstream repo (the Python fork builds the fast tokenizer at
runtime via transformers). The Rust engine's tokenizer loader (mlx_gen::TextTokenizer, consumed by
the qwen-image provider's load_tokenizer) reads the HF tokenizers fast serialization, so it
needs a tokenizer.json.
This repo hosts that derived tokenizer.json so SceneWorks model-install can overlay it onto the
upstream Qwen-Image snapshot (instead of running a Python vocab.json+merges.txt→fast conversion at
install time on every machine — the desktop Mac bundle ships no Python). See SceneWorks sc-6570; this
mirrors the Kolors fast-tokenizer overlay
(sc-4764).
Note:
Qwen/Qwen-Image-Edit-2511already ships its owntokenizer.jsonupstream, so only the base text-to-imageQwen/Qwen-Imagerepo needs this overlay.
Materialized by tools/build_qwen_tokenizer.py (mlx-gen): loads the Qwen2 tokenizer with
transformers.AutoTokenizer.from_pretrained (the fast path) and writes backend_tokenizer.save(...).
The result is the byte-identical fast tokenizer the fork builds at runtime — same vocab, merges,
NFC + ByteLevel pipeline, and special tokens.
Validation: fast-tokenizer ids == the fork's runtime transformers tokenizer across an
EN + EN-long + CN + mixed CN/EN/numeric/punct + empty(negative-prompt) battery — 0 mismatches.
vocab_size 151665, pad token id 151643 (<|endoftext|>).
tokenizer.json — the derived fast tokenizer (the file the Rust engine needs).vocab.json, merges.txt, tokenizer_config.json, added_tokens.json, special_tokens_map.json —
the upstream slow-tokenizer source files (provenance / reproducibility).Derived from the Qwen2 tokenizer shipped with Qwen/Qwen-Image (Apache-2.0). This repo redistributes only the tokenizer (no model weights) for engine interoperability.