Downloads · 30 days
191
43% of all-time downloads
txgsync/Maple-Preview-BF16-MLX
Maple-Preview-BF16-MLX is a text generation model from txgsync. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
This repository contains the full-precision BF16 MLX conversion of deepgrove/maple-preview.
Downloads · 30 days
191
43% of all-time downloads
All-time downloads
447
Public
Parameters
20.2B
40.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors40.4 GB · 100%
From the Hugging Face model README
This repository contains the full-precision BF16 MLX conversion of deepgrove/maple-preview.
maple.py.trust_remote_code=True). In oMLX, enable Trust Remote Code for this model.This is an MLX conversion for local inference on Apple Silicon. Please follow the base model's MIT license and usage terms.
Maple is a reasoning-heavy model and may spend a substantial part of its response budget thinking. For the OpenAI-compatible API or oMLX UI, start with:
temperature: 1.0
top_p: 0.95
top_k: 40
min_p: 0.05
repetition_penalty: 1.0
max_tokens: 8192 or higher
max context: 131072 tokens (native model limit)
These sampler values match DeepGrove's Maple llama.cpp setup. The model declares a native 131,072-token context window and does not require RoPE/YARN scaling for that window. Actual usable context may be lower on systems constrained by KV-cache memory; do not assume that extending beyond 131,072 tokens is supported.
DeepGrove · 2026
Today we introduce Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM. Maple-Preview has SOTA reasoning for its weight class and is even competitive with larger models. It solves IMO-level problems and runs at 200+ tokens/sec on a Mac mini M4, 5–16× faster than efficient models like Gemma 4, Qwen3.5, and gpt-oss.

[!NOTE] The included Transformers implementation depends on Triton and FlashAttention and is intended for a compatible CUDA environment. The reported Apple Silicon result uses a separate on-device runtime.
Maple-Preview is a 20B-A1B reasoning model designed from the start for efficient on-device inference. It utilizes a 24-layer, 256-expert (8 active) configuration with 3:1 SWA-512:GA attention.
On benchmarks, Maple-Preview sets a new point on the Pareto frontier for both memory-to-performance and speed-to-performance, demonstrating its strong reasoning capabilities. However, we note that this preview is focused primarily on raw reasoning and, as such, may underperform on agentic benchmarks. We intend to continue improving general performance through extended training before Maple's full release.

Capability comparison using the dense output head across LCBv6, AIME 2026, HMMT 2026, and GPQA-D.
This preview received minimal post-training for agentic tasks and only small-scale general reinforcement learning.
Maple-Preview is released under the MIT License.