Downloads · 30 days
171
16% of all-time downloads
FastFlowLM/LFM2.5-1.2B-Thinking-NPU2
LFM2.5-1.2B-Thinking-NPU2 is a text generation model from FastFlowLM. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
<div align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inlin…
Downloads · 30 days
171
16% of all-time downloads
All-time downloads
1.1K
Public
Repo size
2 GB
Likes
0
Public
Click a slice to open those files.
.q4nx1 GB · 100%
From the Hugging Face model README
LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Find more information about LFM2.5 in our blog post.
| Model | Parameters | Description |
|---|---|---|
| LFM2.5-1.2B-Base | 1.2B | Pre-trained base model for fine-tuning |
| LFM2.5-1.2B-Instruct | 1.2B | General-purpose instruction-tuned model |
| LFM2.5-1.2B-Thinking | 1.2B | General-purpose reasoning model |
| LFM2.5-1.2B-JP | 1.2B | Japanese-optimized chat model |
| LFM2.5-VL-1.6B | 1.6B | Vision-language model with fast inference |
| LFM2.5-Audio-1.5B | 1.5B | Audio-language model for speech and text I/O |
LFM2.5-1.2B-Thinking is a general-purpose text-only model with the following features:
temperature: 0.1top_k: 50top_p: 0.1repetition_penalty: 1.05| Model | Description |
|---|---|
| LFM2.5-1.2B-Thinking | Original model checkpoint in native format. Best for fine-tuning or inference with Transformers and vLLM. |
| LFM2.5-1.2B-Thinking-GGUF | Quantized format for llama.cpp and compatible tools. Optimized for CPU inference and local deployment with reduced memory usage. |
| LFM2.5-1.2B-Thinking-ONNX | ONNX Runtime format for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). |
| LFM2.5-1.2B-Thinking-MLX | MLX format for Apple Silicon. Optimized for fast inference on Mac devices using the MLX framework. |
We recommend using it for agentic tasks, data extraction, and RAG. It is not recommended for knowledge-intensive tasks and programming.
LFM2.5 uses a ChatML-like format. See the Chat Template documentation for details. Example:
<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
You can use tokenizer.apply_chat_template() to format your messages automatically.
LFM2.5 supports function calling as follows:
tokenizer.apply_chat_template() function with tools.<|tool_call_start|> and <|tool_call_end|> special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt.See the Tool Use documentation for the full guide. Example:
<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>
LFM2.5 is supported by many inference frameworks. See the Inference documentation for the full list.
| Name | Description | Docs | Notebook |
|---|---|---|---|
| Transformers | Simple inference with direct access to model internals. | <a href="https://docs.liquid.ai/lfm/inference/transformers">Link</a> | <a href="https://colab.research.google.com/drive/1_q3jQ6LtyiuPzFZv7Vw8xSfPU5FwkKZY?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| vLLM | High-throughput production deployments with GPU. | <a href="https://docs.liquid.ai/lfm/inference/vllm">Link</a> | <a href="https://colab.research.google.com/drive/1VfyscuHP8A3we_YpnzuabYJzr5ju0Mit?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| llama.cpp | Cross-platform inference with CPU offloading. | <a href="https://docs.liquid.ai/lfm/inference/llama-cpp">Link</a> | <a href="https://colab.research.google.com/drive/1ohLl3w47OQZA4ELo46i5E4Z6oGWBAyo8?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| MLX | Apple's machine learning framework optimized for Apple Silicon. | <a href="https://docs.liquid.ai/lfm/inference/mlx">Link</a> | — |
| LM Studio | Desktop application for running LLMs locally. | <a href="https://docs.liquid.ai/lfm/inference/lm-studio">Link</a> | — |
Here's a quick start example with Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
model_id = "LiquidAI/LFM2.5-1.2B-Thinking"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
# attn_implementation="flash_attention_2" <- uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
prompt = "What is C. elegans?"
input_ids = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
).to(model.device)
output = model.generate(
input_ids,
do_sample=True,
temperature=0.1,
top_k=50,
top_p=0.1,
repetition_penalty=1.05,
max_new_tokens=512,
streamer=streamer,
)
We recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.
| Name | Description | Docs | Notebook |
|---|---|---|---|
| SFT (Unsloth) | Supervised Fine-Tuning with LoRA using Unsloth. | <a href="https://docs.liquid.ai/lfm/fine-tuning/unsloth">Link</a> | <a href="https://colab.research.google.com/drive/1HROdGaPFt1tATniBcos11-doVaH7kOI3?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (TRL) | Supervised Fine-Tuning with LoRA using TRL. | <a href="https://docs.liquid.ai/lfm/fine-tuning/trl">Link</a> | <a href="https://colab.research.google.com/drive/1j5Hk_SyBb2soUsuhU0eIEA9GwLNRnElF?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| DPO (TRL) | Direct Preference Optimization with LoRA using TRL. | <a href="https://docs.liquid.ai/lfm/fine-tuning/trl">Link</a> | <a href="https://colab.research.google.com/drive/1MQdsPxFHeZweGsNx4RH7Ia8lG8PiGE1t?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
We compared LFM2.5-1.2B-Thinking with relevant sub-2B models on a diverse suite of benchmarks.
| Model | GPQA | MMLU-Pro | IFEval | IFBench | Multi-IF | AIME25 | BFCLv3 |
|---|---|---|---|---|---|---|---|
| LFM2.5-1.2B-Thinking | - | - | - | - | - | - | - |
| LFM2.5-1.2B-Instruct | 38.89 | 44.35 | 86.23 | 47.33 | 60.98 | 14.00 | 49.12 |
| Qwen3-1.7B (Thinking) | - | - | - | - | - | - | - |
| Qwen3-1.7B (Instruct) | 34.85 | 42.91 | 73.68 | 21.33 | 56.48 | 9.33 | 46.30 |
| Granite 4.0-1B | 24.24 | 33.53 | 79.61 | 21.00 | 43.65 | 3.33 | 52.43 |
| Llama 3.2 1B Instruct | 16.57 | 20.80 | 52.37 | 15.93 | 30.16 | 0.33 | 21.44 |
| Gemma 3 1B IT | 24.24 | 14.04 | 63.25 | 20.47 | 44.31 | 1.00 | 16.64 |
GPQA, MMLU-Pro, IFBench, and AIME25 follow ArtificialAnalysis's methodology. For IFEval and Multi-IF, we report the average score across strict and loose prompt and instruction accuracies. For BFCLv3, we report the final weighted average score with a custom Liquid handler to support our tool use template.
LFM2.5-1.2B-Thinking offers extremely fast inference speed on CPUs with a low memory profile compared to similar-sized models.

In addition, we are partnering with AMD, Qualcomm, Nexa AI, and FastFlowLM to bring the LFM2.5 family to NPUs. These optimized models are available through our partners, enabling highly efficient on-device inference.
We report prefill throughput evaluated over a range of prompt lengths.
| Platform / Device | Inference | Framework | Model | 1K Prefill (tok/s) | 4K Prefill (tok/s) | 16K Prefill (tok/s) | Memory |
|---|---|---|---|---|---|---|---|
| AMD Ryzen™ AI 395+ | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1,487 | 2,226 | 1,670 | 1.6 GB (full context) |
| AMD Ryzen™ AI 9 HX 370 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1,487 | 2,226 | 1,670 | 1.6 GB (full context) |
| AMD Ryzen™ AI 7 HX 350 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1,431 | 2,032 | 1,519 | 1.6 GB (full context) |
| AMD Ryzen™ AI 5 HX 340 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1,431 | 2,032 | 1,519 | 1.6 GB (full context) |
| AMD Ryzen™ AI 9 HX 370 | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking | 2,975 | N/A | N/A | 856 MB |
| Qualcomm Snapdragon® X Elite | NPU | NexaML | LFM2.5-1.2B-Thinking | 2,591 | N/A | N/A | 0.9 GB |
| Qualcomm Snapdragon® Gen4 (ROG Phone 9 Pro) | NPU | NexaML | LFM2.5-1.2B-Thinking | 4,391 | N/A | N/A | 0.9 GB |
| Qualcomm Dragonwing IQ9 (IQ-9075, IoT) | NPU | NexaML | LFM2.5-1.2B-Thinking | 2,143 | N/A | N/A | 0.9 GB |
| Qualcomm Snapdragon® Gen4 (Galaxy S25 Ultra) | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking | 335 | N/A | N/A | 719 MB |
The reported results correspond to decoding 100 tokens at different context lengths.
| Platform / Device | Inference | Framework | Model | Decode @1K (tok/s) | Decode @4K (tok/s) | Decode @16K (tok/s) | Memory |
|---|---|---|---|---|---|---|---|
| AMD Ryzen™ AI 395+ | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 60 | 54 | 49 | 1.6 GB (full context) |
| AMD Ryzen™ AI 9 HX 370 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 57 | 54 | 49 | 1.6 GB (full context) |
| AMD Ryzen™ AI 7 HX 350 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 63 | 59 | 52 | 1.6 GB (full context) |
| AMD Ryzen™ AI 5 HX 340 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 63 | 59 | 52 | 1.6 GB (full context) |
| AMD Ryzen™ AI 9 HX 370 | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking | 116 | N/A | N/A | 856 MB |
| Qualcomm Snapdragon® X Elite | NPU | NexaML | LFM2.5-1.2B-Thinking | 63 | N/A | N/A | 0.9 GB |
| Qualcomm Snapdragon® Gen4 (ROG Phone 9 Pro) | NPU | NexaML | LFM2.5-1.2B-Thinking | 82 | N/A | N/A | 0.9 GB |
| Qualcomm Dragonwing IQ9 (IQ-9075, IoT) | NPU | NexaML | LFM2.5-1.2B-Thinking | 53 | N/A | N/A | 0.9 GB |
| Qualcomm Snapdragon® Gen4 (Galaxy S25 Ultra) | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking | 70 | N/A | N/A | 719 MB |
LFM2.5-1.2B-Thinking excels at long-context inference.
On AMD NPUs with FastFlowLM, decoding throughput sustains ~46 tok/s even at the full 32K context, indicating robust long-context scalability.
See detailed benchmark results (up to full context length) here.
These capabilities unlock new deployment scenarios across various devices, including vehicles, mobile devices, laptops, IoT devices, and embedded systems.
For enterprise solutions and edge deployment, contact [email protected].
@article{liquidai2025lfm2,
title={LFM2 Technical Report},
author={Liquid AI},
journal={arXiv preprint arXiv:2511.23404},
year={2025}
}