Downloads · 30 days
181
100% of all-time downloads
inclusionAI/Qwen3.6-27B-singprobe
Qwen3.6-27B-singprobe is a machine learning model from inclusionAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
SingProbe is an intrinsic streaming guardrail built on Qwen/Qwen3.6-27B. Rather than running a separate safety model, this lightweight probe reuses the base model's hidden states during generation to score, at every t…
Downloads · 30 days
181
100% of all-time downloads
All-time downloads
181
Public
Parameters
10.1M
20.2 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors20.2 MB · 100%
From the Hugging Face model README
SingProbe is an intrinsic streaming guardrail built on Qwen/Qwen3.6-27B. Rather than running a separate safety model, this lightweight probe reuses the base model's hidden states during generation to score, at every token, query intent, response unsafety, and hallucination risk. It adds less than 0.5% decode-time overhead.
| Base model | Probe parameters | Tapped layers | Outputs |
|---|---|---|---|
inclusionAI/Qwen3.6-27B-singprobe | 10.1M | [20, 41, 62] | 8 intents + unsafe + hallucination |
See the technical report for methodology and complete results. Training codes are available at inclusionAI/SingProbe.
Higher is better for every metric. Results are averages over the benchmark suites specified below.
| Task | Metric | Qwen3.6-27B-singprobe | Reference baseline |
|---|---|---|---|
| Query intent classification (6 benchmarks) | F1 | 0.8718 | YuFeng-XGuard-Reason-8B: 0.8714 |
| Response safety classification (8 benchmarks) | F1 | 0.8692 | Qwen3Guard-Gen-8B-strict: 0.8604 |
| Streaming safety (3 benchmarks) | R-AUC / T-AUC | 0.9864 / 0.9270 | Qwen3Guard-Stream-8B-strict: 0.9640 / 0.8893 |
| Hallucination detection (6 benchmarks) | AUC | 0.8094 | DRIFT: 0.8000 |
| Deployment characteristic | Result |
|---|---|
| Benign-response false-positive rate | 0.02% average across 5 datasets |
| Decode overhead | < 0.5% |
SingProbe is supported through the SGLang integration branch or vLLM integration branch. Load the probe by its Hugging Face ID at server launch:
python -m sglang.launch_server \
--model-path Qwen/Qwen3.6-27B \
--probe-ckpt inclusionAI/Qwen3.6-27B-singprobe \
--port 30000
The integrations return one score dictionary per generated token (label_0–label_9). Use the exact base-model/probe pair: Qwen/Qwen3.6-27B with this checkpoint.
@article{singteam2026singprobe,
title = {SingProbe Technical Report},
author = {Sing Team},
journal = {arXiv preprint arXiv:2608.30703},
year = {2026},
}