Downloads · 30 days
0
lyte-codes/apex-vq-1024
apex-vq-1024 is a machine learning model from lyte-codes. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for numpy. The card lists the license as apache-2.0.
An image tokenizer. Turns a picture into 1,024 discrete tokens — apexes — and back into an approximation of it. Model 263701, the first of the Apex family.
Downloads · 30 days
0
Access
Public
Updated Sep 9, 2026
Repo size
689 KB
Likes
0
Public
Click a slice to open those files.
.npz689 KB · 98%
From the Hugging Face model README
An image tokenizer. Turns a picture into 1,024 discrete tokens — apexes —
and back into an approximation of it. Model 263701, the first of the Apex
family.
Named for the solar apex, the point on the sky the Sun is travelling toward. The codebook is a fixed set of reference directions and every patch is matched to the nearest one, which is the same idea.
from imagetok import PatchVQTokenizer
from PIL import Image
tok = PatchVQTokenizer.load("imagetok_1024.npz")
ids = tok.encode(Image.open("photo.jpg")) # 1024 ints, each 0..1023
back = tok.decode(ids) # a PIL image
A text tokenizer is exactly invertible — byte-level BPE returns the bytes it was given, and anything less is broken. An image tokenizer cannot do that and nothing will make it: a photograph is millions of continuous values and a token is one integer out of a thousand. The mapping discards information by design.
So fidelity is reported as a quantity, not a pass mark. There is no pass mark, and whether coarse is acceptable depends on what the tokens are for.
| tokens per image | 1,024 (32×32 grid of 8×8 patches, 256×256 input) |
| bits per token | 10 |
| size as tokens | 1,280 bytes |
| size as raw pixels | 196,608 bytes |
| compression | 154× smaller |
| reconstruction | 23.4 dB, mean absolute error 13.0/255 |
| codebook used | 970 of 1,024 (95%), usage perplexity 462 |
Quadrupling the codebook from 256 to 1024 buys 0.9 dB (22.5 → 23.4). Halving the patch size buys 2.3 dB and costs four times the tokens:
| patch | tokens/image | reconstruction |
|---|---|---|
| 16×16 | 256 | 21.6 dB |
| 8×8 | 1,024 | 23.9 dB |
| 4×4 | 4,096 | 26.6 dB |
So the loss is dominated by the patch grid, not the codebook. An 8×8 patch cannot represent an edge crossing it at an angle however many codes exist, which is the visible failure — straight edges go blocky while colour and layout survive.
That also bounds anything built on top: a model that generates these tokens can never produce an image better than this tokenizer can reconstruct. If you are using this as the front end of a generative or editing model, 23.4 dB is your ceiling, and a learned convolutional VQ-VAE at the same token count would raise it — the encoder can place information anywhere in the patch instead of matching against a fixed dictionary.
k-means over patches drawn from the training corpus, k-means++ seeded so that a corpus of mostly-sky does not put every code in the sky. Nearest code wins and its index is the token. This is what a VQ-VAE does with a neural encoder in front; without one it stays inspectable — every token is a patch you can look at, so codebook collapse is visible rather than theoretical.
Trained on 272 CC-licensed photographs from Wikimedia Commons, held out 48 for the numbers above. A larger and more varied corpus would move these figures; they are reported on what was actually used.
Apache 2.0. imagetok.py and metric.py are included so the numbers can be
reproduced rather than taken on trust.