Downloads · 30 days
772
3% of all-time downloads
poolside/Laguna-S-2.1-DFlash
Laguna-S-2.1-DFlash is a machine learning model from poolside. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
<p align="center" <img alt="poolside-banner" src="https://poolside.ai/assets/laguna/laguna-s-2-1-banner.svg" width="800px" </p
Downloads · 30 days
772
3% of all-time downloads
All-time downloads
25.5K
Public
Parameters
1.1B
11.2 GB on disk
Likes
26
Public
Click a slice to open those files.
.safetensors2.2 GB · 100%
From the Hugging Face model README
DFlash speculator (draft model) for Laguna S 2.1, bf16.
DFlashLagunaForCausalLM (6 sliding-attention layers, block_size 16).draft_vocab_size == vocab_size (no d2t/t2d).laguna_dflash) and TRT-LLM (pytorch DFlash backend) as the draft model in a speculative config.vllm serve --model poolside/Laguna-S-2.1 \
--speculative-config '{"model":"poolside/Laguna-S-2.1-DFlash","num_speculative_tokens":7,"method":"dflash"}' \
...
python -m sglang.launch_server --model-path poolside/Laguna-S-2.1 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path poolside/Laguna-S-2.1-DFlash \
...
A llama.cpp conversion of this draft model (laguna-s-2.1-DFlash-BF16.gguf) is
published in
poolside/Laguna-S-2.1-GGUF
alongside the target GGUFs. It embeds the target tokenizer and the DFlash metadata
(dflash.decoder_arch = laguna, capture layers, block size), so it works directly
as the -md draft model:
git clone --branch laguna https://github.com/poolsideai/llama.cpp
cd llama.cpp && cmake -B build && cmake --build build -j
./build/bin/llama-server \
-m laguna-s-2.1-Q4_K_M.gguf \
-md laguna-s-2.1-DFlash-BF16.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 \
-fa on --jinja --port 8000
--spec-draft-n-max is clamped to the trained block size (15 draft tokens + 1).
[!NOTE] Requires Poolside's llama.cpp fork, branch
laguna. Upstream llama.cpp ships the generic DFlash framework but not the Laguna decoder contract this draft model needs, and upstream PR ggml-org/llama.cpp#25165 covers the target architecture only.