Downloads · 30 days
40K
97% of all-time downloads
XHToken/Spark-X2.5-4B
Spark-X2.5-4B is a text generation model from XHToken. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
40K
97% of all-time downloads
All-time downloads
41.2K
Public
Parameters
4.1B
8.2 GB on disk
Likes
1.4K
Trending 37
Click a slice to open those files.
.safetensors8.2 GB · 100%
From the Hugging Face model README
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages.
Technical Highlights:
For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture. This approach leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size—thereby improving its practicality and effectiveness across real-world deployment scenarios.
<p align="center"> <img src="./images/spark25-hybrid-architecture-light.png" alt="Spark-X2.5 hybrid architecture"> </p>Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies are conducted to determine an effective balance among mathematics, logic, code, and other high-value domains. This enables the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.
Post-training begins with supervised fine-tuning on a carefully curated corpus. This stage establishes robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning. We subsequently apply large-scale reinforcement learning across several capability domains, including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. This process yields a set of domain-specialized teacher policies, whose complementary strengths are consolidated into a single deployable model through MOPD.
<p align="center"> <img src="./images/post_training_pipeline.svg" alt="Spark-X2.5 hybrid architecture"> </p>We evaluate our models and compare them with leading on-device models of similar size across a broad range of tasks, including agent, code, math, general and knowledge.
<div style="overflow-x: auto; width: 100%;"> <table style="width: 100%; min-width: 1080px; border-collapse: collapse; text-align: center;"> <colgroup> <col style="width: 190px;"> <col span="8" style="width: 110px;"> </colgroup> <thead> <tr> <th align="center">Benchmark</th> <th align="center"><span style="white-space: nowrap;">Spark‑X2.5‑4B</span></th> <th align="center"><span style="white-space: nowrap;">Spark‑X2.5‑1.7B</span></th> <th align="center"><span style="white-space: nowrap;">Qwen3.5‑9B</span></th> <th align="center"><span style="white-space: nowrap;">Qwen3.5‑4B</span></th> <th align="center"><span style="white-space: nowrap;">Qwen3.5‑2B</span></th> <th align="center"><span style="white-space: nowrap;">Gemma4‑12B</span></th> <th align="center"><span style="white-space: nowrap;">Gemma4‑E4B</span></th> <th align="center"><span style="white-space: nowrap;">Gemma4‑E2B</span></th> </tr> </thead> <tbody> <tr><th colspan="9" align="left">Agent</th></tr> <tr><td align="center">BFCL‑V4</td><td align="center">65.1</td><td align="center">46.9</td><td align="center"><strong>66.1*</strong></td><td align="center">50.3*</td><td align="center">43.6*</td><td align="center">37.4</td><td align="center">36.9</td><td align="center">30.2</td></tr> <tr><td align="center">τ²‑bench</td><td align="center">75.1</td><td align="center">65.3</td><td align="center">79.1*</td><td align="center"><strong>79.9*</strong></td><td align="center">48.8*</td><td align="center">69.0*</td><td align="center">42.2*</td><td align="center">24.5*</td></tr> <tr><td align="center">τ³‑bench</td><td align="center"><strong>30.4</strong></td><td align="center">20.1</td><td align="center">9.3</td><td align="center">6.7</td><td align="center">4.1</td><td align="center">13.3</td><td align="center">10.1</td><td align="center">8.8</td></tr> <tr><td align="center">MCP‑Atlas</td><td align="center"><strong>54.6</strong></td><td align="center">23.4</td><td align="center">47.4*</td><td align="center">40.8*</td><td align="center">14.8</td><td align="center">30.5*</td><td align="center">15.0*</td><td align="center">12.6</td></tr> <tr><td align="center">MCP‑Mark</td><td align="center"><strong>14.2</strong></td><td align="center">2.3</td><td align="center">13.4</td><td align="center">12.5</td><td align="center">–</td><td align="center">–</td><td align="center">–</td><td align="center">–</td></tr> <tr><td align="center"><span style="white-space: nowrap;">Workspace Bench</span></td><td align="center"><strong>31.2</strong></td><td align="center">18.9</td><td align="center">25.5</td><td align="center">21.3</td><td align="center">7.7</td><td align="center">–</td><td align="center">–</td><td align="center">–</td></tr> <tr><td align="center">VitaBench2.0</td><td align="center"><strong>25.2</strong></td><td align="center">8.3</td><td align="center">15.6</td><td align="center">18.2</td><td align="center">5.2</td><td align="center">12.4</td><td align="center">4.8</td><td align="center">4.4</td></tr> <tr><td align="center">BrowseComp</td><td align="center"><strong>40.9</strong></td><td align="center">29.7</td><td align="center">8.3</td><td align="center">14.3</td><td align="center">3.1</td><td align="center">10.0</td><td align="center">8.3</td><td align="center">3.7</td></tr> <tr><th colspan="9" align="left">Code</th></tr> <tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Pro</span></td><td align="center"><strong>44.4</strong></td><td align="center">10.4</td><td align="center">33.8*</td><td align="center">29.4*</td><td align="center">1.9</td><td align="center">21.9*</td><td align="center">4.0*</td><td align="center">–</td></tr> <tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Verified</span></td><td align="center">41.6</td><td align="center">28.3</td><td align="center"><strong>53.1*</strong></td><td align="center">38.8*</td><td align="center">6.8</td><td align="center">44.2*</td><td align="center">14.0*</td><td align="center">–</td></tr> <tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Multilingual</span></td><td align="center"><strong>53.3</strong></td><td align="center">23.3</td><td align="center">43.3</td><td align="center">27.7</td><td align="center">5.0</td><td align="center">32.5*</td><td align="center">–</td><td align="center">–</td></tr> <tr><td align="center">SciCode</td><td align="center">34.7</td><td align="center">18.2</td><td align="center">32.7*</td><td align="center">24.0</td><td align="center">6.0</td><td align="center"><strong>39.8</strong></td><td align="center">27.5</td><td align="center">20.5</td></tr> <tr><th colspan="9" align="left">Math</th></tr> <tr><td align="center"><span style="white-space: nowrap;">Gaokao 2026</span></td><td align="center">133.4</td><td align="center">114.8</td><td align="center"><strong>135.5</strong></td><td align="center">130.3</td><td align="center">94.0</td><td align="center">130.6</td><td align="center">102.4</td><td align="center">81.8</td></tr> <tr><td align="center"><span style="white-space: nowrap;">AIME 2026</span></td><td align="center"><strong>90.7</strong></td><td align="center">69.4</td><td align="center">88.2</td><td align="center">83.0</td><td align="center">30.8</td><td align="center">82.1*</td><td align="center">42.5*</td><td align="center">37.5*</td></tr> <tr><td align="center"><span style="white-space: nowrap;">HMMT Feb 2026</span></td><td align="center"><strong>81.2</strong></td><td align="center">48.4</td><td align="center">70.8</td><td align="center">69.7</td><td align="center">21.5</td><td align="center">65.6</td><td align="center">34.2</td><td align="center">20.5</td></tr> <tr><td align="center"><span style="white-space: nowrap;">IMO‑AnswerBench</span></td><td align="center"><strong>74.2</strong></td><td align="center">45.4</td><td align="center">69.8</td><td align="center">68.5</td><td align="center">–</td><td align="center">57.2</td><td align="center">26.9</td><td align="center">22.6</td></tr> <tr><th colspan="9" align="left">General & Knowledge</th></tr> <tr><td align="center">IFEval</td><td align="center">93.0</td><td align="center">89.5</td><td align="center">91.5*</td><td align="center">89.8*</td><td align="center">78.6*</td><td align="center"><strong>94.8</strong></td><td align="center">45.3</td><td align="center">34.8</td></tr> <tr><td align="center">IFBench</td><td align="center"><strong>75.0</strong></td><td align="center">66.3</td><td align="center">64.5</td><td align="center">59.2</td><td align="center">41.3*</td><td align="center">73.5*</td><td align="center">44.0*</td><td align="center">22.7</td></tr> <tr><td align="center">AA‑LCR</td><td align="center">56.3</td><td align="center">24.3</td><td align="center"><strong>63.0*</strong></td><td align="center">57.0*</td><td align="center">25.6*</td><td align="center">55.3*</td><td align="center">34.7</td><td align="center">18.3</td></tr> <tr><td align="center">HLE</td><td align="center">12.3</td><td align="center">6.3</td><td align="center"><strong>14.3</strong></td><td align="center">8.6</td><td align="center">2.1</td><td align="center">13.1</td><td align="center">3.9</td><td align="center">2.5</td></tr> <tr><td align="center">GPQA</td><td align="center">67.4</td><td align="center">43.8</td><td align="center"><strong>77.2</strong></td><td align="center">67.2</td><td align="center">44.6</td><td align="center">72.8</td><td align="center">54.5</td><td align="center">43.8</td></tr> </tbody> </table> </div>The examples below serve a local Spark-X2.5-4B checkpoint. Set MODEL_PATH to its absolute path before starting a container:
export MODEL_PATH=/absolute/path/to/Spark-X2.5-4B
Use the pre-built image that tracks the Spark-X2.5 runtime:
docker pull lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1
# A3 daily build
export SGLANG_IMAGE=quay.io/ascend/sglang:main-cann9.0.0-a3
# A2 daily build (use this instead on A2 hardware)
export SGLANG_IMAGE=quay.io/ascend/sglang:main-cann9.0.0-910b
docker pull "$SGLANG_IMAGE"
The following commands start an OpenAI-compatible API server configured for a maximum context length of 1,048,576 tokens. This setting requires sufficient device memory; reduce --context-length when necessary.
docker run --rm -it \
--gpus '"device=0"' \
--ipc=host \
-p 30000:30000 \
-v "$MODEL_PATH:/root/Spark-X2.5-4B:ro" \
lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
python -m sglang.launch_server \
--model-path /root/Spark-X2.5-4B \
--served-model-name spark2.5 \
--tool-call-parser spark25 \
--reasoning-parser qwen3 \
--tp-size 1 \
--mem-fraction-static 0.8 \
--context-length 1048576 \
--chat-template /root/Spark-X2.5-4B/chat_template.jinja \
--host 0.0.0.0 \
--port 30000
docker run -it --rm -e ASCEND_USE_FIA=1 --network=host --ipc=host --shm-size=16g \
--device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \
--device=/dev/davinci4 --device=/dev/davinci5 --device=/dev/davinci6 --device=/dev/davinci7 \
--device=/dev/davinci8 --device=/dev/davinci9 --device=/dev/davinci10 --device=/dev/davinci11 \
--device=/dev/davinci12 --device=/dev/davinci13 --device=/dev/davinci14 --device=/dev/davinci15 \
--device=/dev/davinci_manager \
--device=/dev/devmm_svm \
--device=/dev/hisi_hdc \
--volume /usr/local/sbin:/usr/local/sbin \
--volume /usr/local/Ascend/driver:/usr/local/Ascend/driver \
--volume /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \
--volume /etc/ascend_install.info:/etc/ascend_install.info \
--volume /var/queue_schedule:/var/queue_schedule \
--volume ~/.cache/:/root/.cache/ \
--volume "$MODEL_PATH:/root/Spark-X2.5-4B:ro" \
--entrypoint=python \
"$SGLANG_IMAGE" \
-m sglang.launch_server \
--model-path /root/Spark-X2.5-4B \
--served-model-name spark2.5 \
--tool-call-parser spark25 \
--reasoning-parser qwen3 \
--tp-size 1 \
--mem-fraction-static 0.8 \
--context-length 1048576 \
--chat-template /root/Spark-X2.5-4B/chat_template.jinja \
--host 0.0.0.0 \
--port 30000
Thinking is enabled by default by both the chat template and the Qwen3 reasoning parser. To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false}.
curl -s http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "spark2.5",
"messages": [
{
"role": "user",
"content": "What is the capital of Anhui Province?"
}
],
"max_tokens": 131072,
"temperature": 1,
"top_k": -1,
"top_p": 0.95,
"repetition_penalty": 1,
"presence_penalty": 0,
"frequency_penalty": 0
}'
vLLM provides an official Docker image for NVIDIA GPU deployment:
docker run --rm --gpus all \
--ipc=host \
-p 30000:30000 \
-v "$MODEL_PATH:/models/Spark-X2.5-4B:ro" \
vllm/vllm-openai:latest \
--model /models/Spark-X2.5-4B \
--port 30000 \
--trust-remote-code \
--served-model-name spark25 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.7 \
--enable-prefix-caching \
--chat-template /models/Spark-X2.5-4B/chat_template.jinja
For Ascend NPUs, choose an official image for the fastest setup.
export IMAGE=quay.io/ascend/vllm-ascend:nightly-main
docker pull "$IMAGE"
export DEVICE=/dev/davinci0
export MODEL_CACHE="${HOME}/.cache"
mkdir -p "$MODEL_CACHE"
docker run --rm \
--name vllm-ascend \
--shm-size=1g \
--device "$DEVICE" \
--device /dev/davinci_manager \
--device /dev/devmm_svm \
--device /dev/hisi_hdc \
-v /usr/local/dcmi:/usr/local/dcmi \
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
-v /etc/ascend_install.info:/etc/ascend_install.info \
-v "$MODEL_CACHE:/root/.cache" \
-p 8000:8000 \
-it "$IMAGE" bash
export IMAGE=quay.io/ascend/vllm-ascend:nightly-main-a3
docker pull "$IMAGE"
export DEVICE0=/dev/davinci0
export DEVICE1=/dev/davinci1
export MODEL_CACHE="${HOME}/.cache"
mkdir -p "$MODEL_CACHE"
docker run --rm \
--name vllm-ascend \
--shm-size=1g \
--device "$DEVICE0" \
--device "$DEVICE1" \
--device /dev/davinci_manager \
--device /dev/devmm_svm \
--device /dev/hisi_hdc \
-v /usr/local/dcmi:/usr/local/dcmi \
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
-v /etc/ascend_install.info:/etc/ascend_install.info \
-v "$MODEL_CACHE:/root/.cache" \
-p 8000:8000 \
-it "$IMAGE" bash
export IMAGE=quay.io/ascend/vllm-ascend:nightly-main-a5
docker pull "$IMAGE"
export MODEL_CACHE="${HOME}/.cache"
mkdir -p "$MODEL_CACHE"
docker run --rm \
--name vllm-ascend \
--net=host \
--shm-size=1g \
--device /dev/davinci0 \
--device /dev/davinci_manager \
--device /dev/devmm_svm \
--device /dev/hisi_hdc \
-v /usr/local/dcmi:/usr/local/dcmi \
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
-v /etc/ascend_install.info:/etc/ascend_install.info \
-v "$MODEL_CACHE:/root/.cache" \
-it "$IMAGE" bash
Install the Spark plugin inside the container:
pip install uv
uv venv ~/spark2_5
source ~/spark2_5/bin/activate
git clone https://github.com/XHToken/Spark-plugin.git
cd ./Spark-plugin
uv pip install .
vllm serve "/models/Spark-X2.5-4B" \
--port "30000" \
--trust-remote-code \
--served-model-name spark25 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.7 \
--enable-prefix-caching \
--chat-template /models/Spark-X2.5-4B/chat_template.jinja
curl -s http://127.0.0.1:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "spark25",
"messages": [{"role": "user", "content": "What is the capital of Anhui Province?"}],
"temperature": 1.0,
"top_k": -1,
"top_p": 0.95
}'
Spark-MLX-LLM runs the original Spark-X2.5 Hugging Face checkpoints locally. It supports Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux. No GGUF conversion is required.
git clone https://github.com/XHToken/Spark-MLX-LLM.git
cd Spark-MLX-LLM
python3 -m venv .venv
source .venv/bin/activate
# Apple silicon
python -m pip install -e .
# Linux CPU
python -m pip install -e '.[cpu]'
# Linux with CUDA 12
python -m pip install -e '.[cuda12]'
# Linux with CUDA 13
python -m pip install -e '.[cuda13]'
spark-mlx-generate \
--device gpu \
--dtype bfloat16 \
--model XHToken/Spark-X2.5-4B \
--prompt "What is the capital of Anhui Province?" \
--max-tokens 512 \
--temp 0
[!NOTE] Native
spark2_5support requires llama.cppb10828or later.
Run the model in the CLI:
llama cli -m /path/to/Spark-X2.5-4B-GGUF
Launch the OpenAI-compatible API server:
llama serve -m /path/to/Spark-X2.5-4B-GGUF
[!NOTE] Native
spark2_5support requires Ollamav0.34.1or later.
ollama run SparkLLM/Spark-X2.5-4B
[!NOTE] Native
spark2_5support requires runtime version2.34.0or later.
[!NOTE] The latest Unsloth Studio supports native Spark-X2.5 GGUF inference.
We recommend using Llama-Factory to fine-tune the model.
The Spark-X2.5 model series is licensed under the Apache 2.0 License.
If you find our work helpful, feel free to give us a cite.
@misc{sparkx2.5,
title = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
author = {SparkLLM Team},
year = {2026}
}