Downloads · 30 days
157
100% of all-time downloads
inclusionAI/Ling-3.0-tiny-singprobe
Ling-3.0-tiny-singprobe is a machine learning model from inclusionAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
<p align="center"<bEnglish</b | <a href="https://huggingface.co/inclusionAI/Ling-3.0-tiny-singprobe/blob/main/READMECN.md"中文</a</p
Downloads · 30 days
157
100% of all-time downloads
All-time downloads
157
Public
Parameters
3.2M
6.4 MB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors6.4 MB · 100%
From the Hugging Face model README
SingProbe is an intrinsic streaming guardrail built on inclusionAI/Ling-3.0-tiny. Rather than running a separate safety model, this lightweight probe reuses the base model's hidden states during generation to score, at every token, query intent, response unsafety, and hallucination risk. It adds less than 0.5% decode-time overhead.
| Base model | Probe parameters | Tapped layers | Outputs |
|---|---|---|---|
inclusionAI/Ling-3.0-tiny-singprobe | 3.22M | [6, 14, 22] | 8 intents + unsafe + hallucination |
See the technical report for methodology and complete results; implementation details are available at inclusionAI/SingProbe.
Higher is better for every metric. Results are averages over the benchmark suites specified below.
| Task | Metric | Ling-3.0-tiny-singprobe | Reference baseline |
|---|---|---|---|
| Query intent classification (6 benchmarks) | F1 | 0.8561 | Qwen3Guard-Stream-8B-strict: 0.8602 |
| Response safety classification (8 benchmarks) | F1 | 0.8508 | Qwen3Guard-Stream-8B-strict: 0.8486 |
| Streaming safety (3 benchmarks) | R-AUC / T-AUC | 0.9888 / 0.9479 | Qwen3Guard-Stream-8B-strict: 0.9640 / 0.8893 |
| Hallucination detection (6 benchmarks) | AUC | 0.7765 | DRIFT: 0.7408 |
| Deployment characteristic | Result |
|---|---|
| Benign-response false-positive rate | 0.07% average across 5 datasets |
| Online hallucination detection | 0.6807 average AUC under free generation |
| Decode overhead | < 0.5% |
SingProbe is supported through the SGLang integration branch or vLLM integration branch. Load the probe by its Hugging Face ID at server launch:
python -m sglang.launch_server \
--model-path inclusionAI/Ling-3.0-tiny \
--probe-ckpt inclusionAI/Ling-3.0-tiny-singprobe \
--port 30000
The integrations return one score dictionary per generated token (label_0–label_9). They currently support Ling-3.0 (BailingMoeV3ForCausalLM) base models only. Use the exact base-model/probe pair: inclusionAI/Ling-3.0-tiny with this checkpoint.
@article{singteam2026singprobe,
title = {SingProbe Technical Report},
author = {Sing Team},
journal = {arXiv preprint arXiv:2608.30703},
year = {2026},
}