Downloads · 30 days
0
pierjoe/limite-mlx
limite-mlx is a machine learning model from pierjoe. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
MLX implementation of the limite architecture, packaged so that stock mlx-lm can load Limite checkpoints with no patching.
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
12.6 KB
Likes
0
Public
Click a slice to open those files.
.py23.8 KB · 45%
From the Hugging Face model README
MLX implementation of the limite architecture, packaged so that stock mlx-lm
can load Limite checkpoints with no patching.
pip install mlx-lm "limite-mlx @ https://huggingface.co/pierjoe/limite-mlx/resolve/main/limite_mlx-0.1.0-py3-none-any.whl"
mlx_lm.generate --model pierjoe/limite-1b-violetto-mlx-4bit \
--prompt "What is the remainder when 7^2026 is divided by 13?" \
--max-tokens 3000 --temp 0.6 --top-p 0.95
Models that use it are in the collection.
mlx-lm resolves an architecture with a single line:
importlib.import_module(f"mlx_lm.models.{model_type}")
and offers no plugin hook, so the usual advice for an unsupported architecture is
to copy a file into site-packages/mlx_lm/models/ — which every mlx-lm upgrade
deletes. This package appends a one-name finder to sys.meta_path at interpreter
startup instead, so that import call resolves without mlx-lm's package directory
being touched.
The finder is appended, not prepended, so the day mlx-lm ships its own
models/limite.py the official implementation wins and this package goes inert.
Ten deviations from a Llama-style decoder, all pinned in the checkpoint's
config.json and reproduced exactly:
finfo(bfloat16).eps = 0.0078125, not 1e-6)0.1, not head_dim ** -0.5qkv_scale / o_scale folded into the projections in bfloat162 * sigmoid(...), before o_proj23 * sigmoid((raw + 5) / 7.5)Ported from the reference vLLM plugin and checked against an independent NumPy oracle with bit-level bfloat16 emulation:
This architecture is unusually rounding-sensitive: the residual stream reaches
|h| ~ 1.5e5 by layer 47, and a single bfloat16 ULP perturbation at layer 1 moves
the final logits by 1.26.
Apache-2.0, matching the upstream model and reference implementation. The model itself is Paradigma's work; this is the MLX port.