Downloads ยท 30 days
18
2% of all-time downloads
Quatfit/Quatfit-Mini-MTP
Quatfit-Mini-MTP is a text generation model from Quatfit. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
<p align="center" <img src="https://huggingface.co/Quatfit/Quatfit-Mini/resolve/main/banner.png" alt="Quatfit Mini MTP Banner" width="100%" </p
Downloads ยท 30 days
18
2% of all-time downloads
All-time downloads
927
Public
Parameters
78.8M
348 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors316 MB ยท 91%
How the weights are stored.
F3278.5M ยท 100%
From the Hugging Face model README
Quatfit Mini MTP is the standalone multi-token-prediction (MTP) drafter used for speculative decoding with Quatfit Mini. Rather than a separate small language model, it is a lightweight autoregressive head that cross-attends directly to the target model's own KV cache โ it has its own token embedder and a compact 4-layer transformer block (3 local sliding-window layers + 1 global layer), but no separate prefill pass and no independent context of its own to maintain.
This repository packages that drafter as a standalone download so it can be attached to any Quatfit Mini serving deployment (FP32, BF16, FP8, or GGUF) to accelerate decoding.
This is not a general-purpose chat model โ it drafts candidate continuations that the target model verifies. Use it together with Quatfit/Quatfit-Mini, Quatfit/Quatfit-Mini-FP8, or Quatfit/Quatfit-Mini-GGUF.
| Metric | Result |
|---|---|
| Decode throughput vs. standard autoregressive decoding | 2.4ร |
| Draft token acceptance rate vs. a same-size standalone draft-model baseline | +97% relative (โ91% vs. โ46%) |
See Methodology below for exactly what these numbers measure and the conditions they were measured under โ read that section before quoting these figures in your own comparisons.
Most speculative-decoding setups pair a large target model with a small, independently-trained draft model that has to run its own forward pass and maintain its own context. Quatfit Mini MTP takes a different approach:
| Component | Value |
|---|---|
| Type | Autoregressive MTP drafter, cross-attends to target model KV cache |
| Target model | Quatfit Mini (8B, Gemma 4 architecture) |
| Drafter layers | 4 (3 local sliding-window + 1 global) |
| Hidden dimension | 512 |
| Attention heads | 4 |
| Own embedder | Yes โ separate from target model's input embeddings |
| Own KV cache | No โ cross-attends to target model's cache |
| Output head | Top-k clustered projection ($d \times 4{,}096$, not full vocab) |
| Parameters | ~180M |
| Precision | BF16 (FP8 variant available for FP8-served targets) |
vllm serve Quatfit/Quatfit-Mini \
--speculative-model Quatfit/Quatfit-Mini-MTP \
--num-speculative-tokens 5 \
--max-model-len 131072
from vllm import LLM, SamplingParams
llm = LLM(
model="Quatfit/Quatfit-Mini",
speculative_model="Quatfit/Quatfit-Mini-MTP",
num_speculative_tokens=5,
max_model_len=131072,
)
sampling_params = SamplingParams(temperature=0.7, max_tokens=512)
outputs = llm.generate(["Explain the difference between GQA and MHA."], sampling_params)
print(outputs[0].outputs[0].text)
vllm serve Quatfit/Quatfit-Mini-FP8 \
--quantization fp8 \
--speculative-model Quatfit/Quatfit-Mini-MTP \
--num-speculative-tokens 5
./llama-server \
-hf Quatfit/Quatfit-Mini-GGUF:Q4_K_M \
--model-draft Quatfit-Mini-MTP-Q4_K_M.gguf \
--draft-max 5 \
-c 131072
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
target = AutoModelForImageTextToText.from_pretrained(
"Quatfit/Quatfit-Mini", torch_dtype="auto", device_map="auto"
)
drafter = AutoModelForImageTextToText.from_pretrained(
"Quatfit/Quatfit-Mini-MTP", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("Quatfit/Quatfit-Mini")
messages = [{"role": "user", "content": "Write a Python implementation of binary search."}]
inputs = processor.apply_chat_template(messages, tokenize=True, return_tensors="pt").to(target.device)
outputs = target.generate(**inputs, assistant_model=drafter, max_new_tokens=512)
print(processor.decode(outputs[0]))
The headline numbers above were measured as follows โ please validate against your own workload before relying on them for capacity planning.
Decode throughput (2.4ร): measured as end-to-end tokens/second during decode (post-prefill), comparing standard single-token autoregressive decoding against speculative decoding with this drafter, both targeting Quatfit Mini at the same precision and batch size, on a single H100. Throughput gain varies with batch size, sequence length, and task mix โ larger batches and less-predictable generation (e.g. open-ended creative writing vs. structured code) typically see a smaller multiple than the figure quoted.
Draft acceptance rate (+97% relative): measured as the fraction of drafted tokens accepted by the target model's verification step, averaged across a mixed evaluation set of code, chat, reasoning, and tool-use prompts, at draft length 5. The "baseline" is an independently trained standalone draft model of comparable parameter count (~180M) with its own embedder, own KV cache, and no cross-attention into the target model โ i.e., the conventional speculative-decoding setup this drafter is designed to improve on, not a specific named third-party product.
| Quatfit Mini (base) | Quatfit Mini FP8 | Quatfit Mini GGUF | Quatfit Mini MTP (this repo) | |
|---|---|---|---|---|
| Role | Target model, full fidelity | Target model, GPU serving | Target model, local/CPU | Drafter, attaches to any of the above |
| Format | safetensors, FP32 | safetensors, FP8 + BF16 | .gguf | safetensors |
| Runtime | ๐ค Transformers | vLLM, TensorRT-LLM, SGLang | llama.cpp | vLLM, TensorRT-LLM, llama.cpp, Transformers |
| Size | ~32 GB | ~9.5 GB | ~3โ16.8 GB per quant | ~180M params (~360 MB BF16) |
Full architecture details for both the target model and this drafter are documented in the Quatfit Mini Technical Report.
This drafter does not generate final output on its own โ all drafted tokens are verified by the target model before being returned, so speculative decoding does not change what the target model would otherwise produce; it only changes how fast those tokens are produced. Standard responsible-use guidance for Quatfit Mini itself still applies:
@article{quatfitminimtp2026,
title={Quatfit Mini MTP: A Cross-Attending Multi-Token-Prediction Drafter for Speculative Decoding},
author={Quatfit AI Research},
year={2026}
}
Apache License 2.0. See LICENSE for details.