Downloads · 30 days
0
SargeDev/jev-distill-corpus
jev-distill-corpus is a machine learning model from SargeDev. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A LoRA adapter (r=16, α=32, on qproj/vproj) on Qwen/Qwen2.5-0.5B-Instruct, distilled from the Jev typed-judgment API into a compact local judge for agent-memory gating.
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
464 MB
Likes
0
Public
Click a slice to open those files.
.jsonl233 MB · 100%
From the Hugging Face model README
A LoRA adapter (r=16, α=32, on q_proj/v_proj) on Qwen/Qwen2.5-0.5B-Instruct, distilled from the Jev typed-judgment API into a compact local judge for agent-memory gating.
What it does: given a query and a candidate memory passage, outputs P(relevant) as the calibrated yes probability read from the final-token logits of yes vs no. Used to filter which vector-recalled memories get injected into agent context (vector recall → cross-encoder band gate → this judge for gray-band cases).
Training: LoRA, 1 epoch, LR 1e-4, on rows sampled from SargeDev/jev-distill-corpus — paired relevance judgments (Jev graded 0–7 expected-value scores + binary labels) distilled from typed judgment calls by a larger teacher model. Prompt format:
Memory: {passage ≤600 chars}
Query: {query}
Question: Is this memory relevant for answering the query? Answer yes or no with confidence.
A 10,000-row held-out test set (test_10k.jsonl, sampled from SargeDev/jev-distill-corpus v2, seed 42, hard-disjoint from all 142,909 train/dev/test rows by (query, text[:600]) key). Gold labels = 32B teacher scores rescaled /7. All three judges scored with the identical prompt above, yes/no logit softmax.
| Judge | MAE | Pearson r | Binary agree @0.5 | Gray-band MAE (n=1,669) | Latency |
|---|---|---|---|---|---|
| Jev 1.13 (API, OpenRouter) | 0.187 | 0.787 | 84.4% | 0.205 | ~1,007 ms |
| Student B (this adapter, local) | 0.219 | 0.709 | 81.7% | 0.275 | 23 ms (batch-32, RTX 3060) |
| Vanilla Qwen2.5-0.5B-Instruct | 0.498 | −0.005 | 44.4% | 0.329 | 22 ms |
Student B vs its own teacher (Jev 1.13): 86.4% agreement, Pearson r = 0.824, MAE 0.144 across all 10k rows. The 1.1M-parameter LoRA reproduces a production System One model's judgments within ~3 points of binary agreement, at 43× lower latency and zero API cost.
Where they disagree (1,364 items), Jev is right vs gold 60% of the time — a real but modest teacher edge. The gray band (teacher y ∈ [0.3, 0.7]) is where both models live in production (the band gate escalates only these cases); student gray-band MAE 0.275 vs teacher 0.205.
| Eval | Student B | Vanilla 0.5B | Student A |
|---|---|---|---|
| Original test (5,605) | 0.148 / 0.836 / 86.0% | 0.443 / −0.007 / 51.8% | 0.112 / 0.886 / 87.6% |
| Held-out (n=60) | 0.187 / 0.791 / 90.0% | 0.536 / −0.067 / 38.3% | 0.273 / 0.445 / 60.0% |
Qdrant vector recall (25 candidates)
→ bge-reranker-base int8 ONNX cross-encoder (local band gate: keep ≥0.08 / drop ≤0.01)
→ gray band [0.01, 0.08] → Student B (this adapter) judges locally
→ final TOP-K injection (default 8)
Fail-open everywhere: any stage error → keep everything (no forgetting).
Every decision is logged with (vector_score, local_score, student_score, verdict) to a JSONL log — that log is the retraining corpus for the next student iteration (see v2 dataset below).
| Repo | Contents |
|---|---|
| SargeDev/jev-gate-student-b | This adapter (adapter_config.json + adapter_model.safetensors + tokenizer) |
| SargeDev/jev-distill-corpus | 148,160 original |
Apache-2.0 (matches base model).