Downloads · 30 days
53
30% of all-time downloads
vectionlabs/Salience-27B-R4
Salience-27B-R4 is a image-text-to-text model from vectionlabs. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
53
30% of all-time downloads
All-time downloads
179
Public
Parameters
27.8B
55.6 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors55.6 GB · 100%
From the Hugging Face model README
A 27B dense vision-language engineer — the full-capacity tier of the Salience family: every parameter on, every token.
Vection Labs
Weights · Quickstart · Limitations
</div>[!Note] Preview release. This is an early build of the Salience 27B tier. It is stable for daily use, but rough edges are expected — anything you report in the Community tab gets fixed in the stable release. This model's dataset contains reasoning traces from Fable 5, but manually reduced for token efficiency.
Salience 27B is a 27-billion-parameter dense vision-language model built for hard, practical engineering work: writing and debugging real code, repo-scale edits, multi-step terminal agency, and quantitative reasoning — with native vision and a 262K-token context window (extendable to 1M via YaRN).
Where the MoE tiers of the family (Pro, Flash) route a few billion active parameters per token, Salience 27B runs all 27B on every token — maximum per-token capacity, a hybrid linear+full attention stack for long-context speed, and an MTP head for self-speculative decoding.
It is engineered for people who care less about chat pleasantries and more about whether the model can do the thing: ship the function, find the bug, drive the terminal, land the pull request.
transformers-native.| Parameters | 27.2B dense (all active) |
| Modalities | text, image, video -> text |
| Context window | 262,144 tokens native (up to 1,048,576 via YaRN) |
| Attention | hybrid linear + full attention (full every 4th layer) |
| Decoding | MTP head included (self-speculative decoding) |
| Precision | bfloat16 |
| Architecture | Qwen3.6 dense (27B) + native vision encoder |
| License | Apache-2.0 |
| Library | 🤗 transformers (AutoModelForImageTextToText) |
The family: Pro (35B-A3B MoE) · Flash (30B-A3B MoE) · 27B Preview (dense) · Nano (9B dense)
Thinking is on by default: the model reasons inside <think>...</think> before answering,
and serving stacks expose it as reasoning_content. Reasoning is native — you never have to
write think step by step (doing so makes it perform reasoning instead of doing it). Control
depth with the token budget, not the prompt. Pass enable_thinking=False to
apply_chat_template for instant direct answers.
The model emits XML-style tool calls (<tool_call><function=...><parameter=...>), parsed
natively by vLLM / SGLang tool parsers for this model family, and by llama-server --jinja.
Provide tool schemas via the chat template tools argument.
Salience 27B R4 targets software engineering, coding agents, and technical research:
It is not intended for high-stakes decisions without human review, nor as a source of truth for medical, legal, or financial advice.
from transformers import AutoModelForImageTextToText, AutoProcessor
import torch
repo = "vectionlabs/Salience-27B-R4"
proc = AutoProcessor.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(
repo, dtype="auto", device_map="auto"
)
messages = [{
"role": "user",
"content": [{"type": "text", "text": "Implement an LRU cache in Python with O(1) get/put."}],
}]
text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = proc(text=[text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=2048)
print(proc.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])
Requires a recent transformers (>= 5.8). Vision works the same way with
{"type": "image", "image": ...} content items.
This is a dense model, so standard quant intuition applies: Q4_K_M and up hold quality well; use Q5_K_M/Q6_K when VRAM allows. (The MoE tiers of the family need Q5/Q6 minimum — that constraint does not apply here.) Keep the MTP layers if your quant includes them: they enable self-speculative decoding for free extra speed.
--jinja with llama-server (or vLLM/SGLang parsers) so XML tool calls
become proper OpenAI-style tool_calls.<sub>Built on Qwen3.6 (Apache-2.0).</sub>
<div align="center"><sub>© 2026 Vection Labs</sub></div>