Downloads · 30 days
134
11% of all-time downloads
vectionlabs/Salience-27B-R5
Salience-27B-R5 is a image-text-to-text model from vectionlabs. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
134
11% of all-time downloads
All-time downloads
1.2K
Public
Parameters
27.8B
55.6 GB on disk
Likes
33
Public
Click a slice to open those files.
.safetensors55.6 GB · 100%
From the Hugging Face model README
A 27B dense vision-language engineer that stops thinking once it has the answer.
Vection Labs
Weights · Reasoning effort · Quickstart · Limitations
</div>[!Note] R5. Fifth revision of the Salience Ridge 27B tier, rebuilt on the Qwen3.8 architecture. Stable for daily use; rough edges get fixed in the stable release — report them in the Community tab.
Salience 27B is a 27-billion-parameter dense vision-language model built for hard, practical engineering work: writing and debugging real code, repo-scale edits, multi-step terminal agency, and quantitative reasoning — with native vision and 1,048,576 tokens of context.
Where the MoE tiers of the family (Pro, Flash) route a few billion active parameters per token, Salience 27B runs all 27B on every token — maximum per-token capacity, a hybrid linear+full attention stack for long-context speed, and an MTP head for self-speculative decoding.
R5's headline change is reasoning economy. A reasoning model pays for accuracy in tokens, and most of them pay the same price for "what does this flag do" as for "why does this deadlock under load". R5 does not: it reasons hard when the problem needs it and answers directly when it does not — and unlike the stock configuration, that is the default behaviour rather than something you have to ask for.
Thinking is on by default: the model reasons inside <think>...</think> before answering,
and serving stacks expose it as reasoning_content. What R5 changes is how much.
| value | behaviour | use it for |
|---|---|---|
low | keeps the chain short and moves straight to the conclusion | chat, lookups, formatting, refactors |
medium | default — no deliberation instruction; the model decides | everyday engineering work |
xhigh | deliberate at length, validate assumptions, weigh alternatives | hard debugging, architecture, math |
# default: proportional reasoning, nothing to configure
text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
# ask for depth when the problem earns it
text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
reasoning_effort="xhigh")
# skip thinking entirely
text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
enable_thinking=False)
Reasoning is native — you never have to write think step by step. Doing so makes a model of this kind perform reasoning instead of doing it.
transformers-native.| Parameters | 27.8B dense (all active) |
| Modalities | text, image, video -> text |
| Context window | 1,048,576 tokens (YaRN + Dual Chunk Attention) |
| Attention | hybrid linear + full attention (full every 4th layer) |
| Decoding | MTP head included (self-speculative decoding) |
| Precision | bfloat16 |
| Architecture | Qwen3.8 dense (27B) + native vision encoder |
| License | Apache-2.0 |
| Library | 🤗 transformers (AutoModelForImageTextToText) |
The family: Pro (35B-A3B MoE) · Flash (30B-A3B MoE) · 27B R5 (dense) · Nano (9B dense)
The model emits XML-style tool calls (<tool_call><function=...><parameter=...>), parsed
natively by vLLM / SGLang tool parsers for this model family, and by llama-server --jinja.
Provide tool schemas via the chat template tools argument.
Salience 27B R5 targets software engineering, coding agents, and technical research:
It is not intended for high-stakes decisions without human review, nor as a source of truth for medical, legal, or financial advice.
18 safetensors shards, 8 JSON configs, a chat template, this README and a banner.
No pickle files, no .bin, no .py. Safetensors is a flat tensor container with no
mechanism for executing code, which is why it replaced pickle — the file list is on this
page and you do not have to take my word for any of it.
Weights cannot run anything on their own; a harness runs things. If you hand any model shell access inside an agent loop that is your risk surface, and it is identical here and for the stock base model.
from transformers import AutoModelForImageTextToText, AutoProcessor
import torch
repo = "vectionlabs/Salience-27B-R5"
proc = AutoProcessor.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(
repo, dtype="auto", device_map="auto"
)
messages = [{
"role": "user",
"content": [{"type": "text", "text": "Implement an LRU cache in Python with O(1) get/put."}],
}]
text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = proc(text=[text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=2048)
print(proc.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])
Requires a recent transformers (>= 5.8). Vision works the same way with
{"type": "image", "image": ...} content items.
| build | format | notes |
|---|---|---|
| bartowski/vectionlabs_Salience-27B-R5-GGUF | GGUF | full range of sizes |
| mradermacher/Salience-27B-R5-i1-GGUF | GGUF | imatrix |
| mradermacher/Salience-27B-R5-GGUF | GGUF | static |
| McG-221/Salience-27B-R5-mlx-8Bit | MLX | Apple silicon |
Thanks to everyone who built those.
llama-server -m Salience-27B-R5-Q4_K_M.gguf \
--jinja --reasoning-format deepseek \
-c 32768 -ngl 999
--jinja is not optional for agent use: it applies the model's own chat template, which is
what turns XML tool calls into proper OpenAI-style tool_calls — and what makes the reasoning
defaults above take effect. Without it you get malformed calls and stock behaviour.
This is a dense model, so standard quant intuition applies: Q4_K_M and up hold quality well; use Q5_K_M/Q6_K when VRAM allows. (The MoE tiers of the family need Q5/Q6 minimum — that constraint does not apply here.) Keep the MTP layers if your quant includes them: they enable self-speculative decoding for free extra speed.
Ships with YaRN (factor 4.0, original_max_position_embeddings 262144) and a
dual_chunk_attention_config block. Static YaRN taxes short prompts slightly; that is the cost
of having the full window available by default. vLLM and SGLang read the DCA block,
transformers ignores it.
reasoning_effort
instead of prompt scaffolding.--jinja with llama-server (or vLLM/SGLang parsers) so XML tool calls
become proper OpenAI-style tool_calls.None have been run yet. Not withheld — not run.
When they exist they will be against a plain Qwen3.8-27B baseline on the same harness and the same day, because a number without its baseline is not a measurement. Until then, treat everything above as a description of what this model was built to do rather than proof that it does it.
medium reasoning by default means shorter chains on genuinely hard problems than a model
pinned to maximum effort. Pass reasoning_effort="xhigh" when the problem deserves it.<sub>Built on Qwen3.8 (Apache-2.0).</sub> <sub>Build with love by the vectionlabs' team (Apache-2.0).</sub>
<div align="center"><sub>© 2026 Vection Labs</sub></div>