Downloads · 30 days
2K
8% of all-time downloads
poolside/Laguna-XS-2.1-FP8
Laguna-XS-2.1-FP8 is a text generation model from poolside. Use it when you need the model to write or continue text. It is set up for vllm. The card lists the license as openmdw-1.1.
<p align="center" <img alt="poolside-banner" src="https://poolside.ai/assets/laguna/laguna-xs2-1-banner.svg" width="800px" </p
Downloads · 30 days
2K
8% of all-time downloads
All-time downloads
24.1K
Public
Parameters
33.4B
35.3 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors35.3 GB · 100%
How the weights are stored.
F8_E4M331.6B · 94%
From the Hugging Face model README
Laguna XS 2.1-FP8 is a 33B total parameter Mixture-of-Experts model with 3B activated parameters per token designed for agentic coding and long-horizon work on a local machine. It uses Sliding Window Attention with per-head gating in 30 out of 40 layers for fast inference and low KV cache requirements.
[!NOTE] This is the FP8 variant with an FP8-quantized KV cache. The BF16, NVFP4 and INT4 variants are also available on Hugging Face.
| Model | Size (total params.) | SWE-bench Verified | SWE-bench Multilingual | SWE-Bench Pro (Public Dataset) | Terminal-Bench 2.0 |
|---|---|---|---|---|---|
| Laguna XS 2.1 | 33B | 70.9% | 63.1% | 47.6% | 37.5% |
| Laguna XS.2 | 33B | 69.9% | 57.7% | 46.3% | 35.7% |
| Qwen3.6-35B-A3B | 35B | 73.4% | 67.2% | 49.5% | 51.5% |
| North Mini Code | 30B | 67.6% | - | 40.2% | 36.0% |
| MAI-Code-1-Flash | 137B | 71.6% | 65.5% | 51.2% | 54.8% |
| gpt-oss-120B | 120B | - | - | 16.2% | 18.7% |
| Claude Haiku 4.5 | - | 73.3% | - | 39.5% | 29.8% |
| GPT-5.4 Nano | - | - | - | 52.4% | 46.3% |
We used the highest publicly-referenced scores for all comparison models across each benchmark. In all cases these were official scores published in release blog posts or equivalent, with the exception of gpt-oss-120b and Claude Haiku 4.5 where the highest published (verified) scores for SWE-Bench Pro and Terminal-Bench 2.0 are from their respective official leaderboards.
<details> <summary>Expand for benchmarking methodology</summary>All benchmarking for Laguna XS 2.1 was completed using Laude Institute’s Harbor Framework with our agent harness, with a maximum of 500 steps and sandboxed execution. The same sampling parameters were used for all Laguna XS 2.1 benchmarking: temperature=1.0, top_k=20 and top_p=1, with thinking mode enabled and a context length of 256K tokens. All tasks were run in their own sandbox using 8 GB RAM/2 CPUs, with the exception of Terminal-Bench 2.0, which used 48 GB RAM/32 CPUs.
Some base task images and verifiers were patched to fix infrastructure reliability issues inherent in task setup, such as rate limits on third-party dependencies in external registries used by the verifier. All four agentic benchmarks were run with patched images. We also ran a reward-hack judge post-hoc on Laguna XS 2.1 evaluation runs and did not find significant reward hacking after joint judge review and manual review.
Laguna XS 2.1-FP8 was evaluated against the unquantized BF16 checkpoint on the same four agentic benchmarks. Scores are mean pass@1; the ± figure is the run-to-run variation and RH is the share of runs flagged by our post-hoc reward-hack judge (see the benchmarking methodology note above).
| Benchmark | Laguna XS 2.1 (BF16) | Laguna XS 2.1-FP8 |
|---|---|---|
| SWE-bench Verified | 70.85% ± 0.85% (6.51% RH) | 70.75% ± 0.85% (6.22% RH) |
| SWE-bench Multilingual | 63.17% ± 1.42% (7.25% RH) | 64.92% ± 1.42% (6.80% RH) |
| SWE-Bench Pro (Public Dataset) | 47.61% ± 0.96% (2.40% RH) | 48.02% ± 0.89% (2.67% RH) |
| Terminal-Bench 2.0 | 37.53% ± 2.81% (2.12% RH) | 40.22% ± 2.70% (3.02% RH) |
Laguna XS 2.1-FP8 has launch-day support in vLLM, SGLang and Transformers, and TRT-LLM thanks to the support of the team at NVIDIA.
[!NOTE] For complete usage instructions, see the main Laguna XS 2.1 model card.
Laguna XS 2.1-FP8 is supported in vLLM, SGLang and Transformers, and TRT-LLM thanks to the support of the team at NVIDIA. Use Laguna XS 2.1 with Ollama (with MLX support) or Llama.cpp (BF16 and Q4_K_M only) for the best results on your local machine.
The full vLLM recipe is on the main Laguna XS 2.1 model card and on the vLLM recipes page. Quantization is detected automatically from quantization_config in this checkpoint, so the same command works with poolside/Laguna-XS-2.1-FP8 substituted for the model ID. Set VLLM_BLOCKSCALE_FP8_GEMM_FLASHINFER=0 when serving with vLLM.
[!IMPORTANT] Tool calling requires a vLLM build with vllm-project/vllm#47311; on older builds pass
--tool-call-parser glm47(the GLM 4.7 parser) instead ofpoolside_v1.
[!NOTE] The FP8-quantized KV cache requires vLLM >= 0.22.0. Earlier versions produce scrambled output on non-Hopper GPUs because of a per-layer attention-head count bug, fixed in vllm#42650. On older vLLM, disable the FP8 KV cache by adding
--kv-cache-dtype-skip-layers $(seq 0 39).
Laguna XS 2.1 is supported in SGLang via sgl-project/sglang#24204. Quantization is detected automatically from quantization_config, so no extra flags are required. See the SGLang cookbook entry and the main Laguna XS 2.1 model card for a serving recipe.
The full Transformers recipe is on the main Laguna XS 2.1 model card. Substitute poolside/Laguna-XS-2.1-FP8 for the model ID; quantization is detected automatically from quantization_config.
Laguna XS 2.1 support ships in TensorRT-LLM >=1.3.0rc16; see the install recipe on the main Laguna XS 2.1 model card. Substitute poolside/Laguna-XS-2.1-FP8 for the model ID; quantization is detected automatically from quantization_config, no extra flags required.
from tensorrt_llm import LLM
llm = LLM(model="poolside/Laguna-XS-2.1-FP8", trust_remote_code=True)
Available on the Ollama library.
[!NOTE] macOS (Metal) users: Chat (
ollama run//api/chat) works as expected on Linux/CUDA. On macOS/Metal it may currently return empty output; the root cause is not yet fully understood and we're investigating it with the Ollama team. On a Mac, use a Linux/CUDA host, or the/api/generateendpoint with"raw": true.
Laguna XS 2.1-FP8 uses the same reasoning controls (interleaved thinking, preserved reasoning, and the enable_thinking flag) as the base model. See the Controlling reasoning section of the main Laguna XS 2.1 model card.
This model is licensed under the OpenMDW-1.1 License.
Laguna XS 2.1-FP8 is designed for software engineering and agentic coding use cases, and you are responsible for confirming that it is appropriate for your intended application. Laguna XS 2.1-FP8 is subject to the OpenMDW-1.1 License, and should be used consistently with Poolside's Acceptable Use Policy. We advise against circumventing Laguna XS 2.1-FP8 safety guardrails without implementing substantially equivalent mitigations appropriate for your use case.
Please report security vulnerabilities or safety concerns to [email protected].