Downloads · 30 days
11K
100% of all-time downloads
alibiserikbay/JevK5
JevK5 is a text generation model from alibiserikbay. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
JevK5 is an independent, Apache-2.0 open-source alternative to TypeSafe's Jev for typed decisions. It reads a state and a yes/no (noul), choice, or score question and returns a probability for every option in one forw…
Downloads · 30 days
11K
100% of all-time downloads
All-time downloads
11K
Public
Parameters
4.2B
25.3 GB on disk
Likes
20
Trending 3
Click a slice to open those files.
.safetensors8.4 GB · 100%
From the Hugging Face model README
JevK5 is an independent, Apache-2.0 open-source alternative to TypeSafe's Jev for typed
decisions. It reads a state and a yes/no (noul), choice, or score question and returns a
probability for every option in one forward pass, with zero generated tokens. The
open weights can be self-hosted using the
JevK5 runtime, which serves a TypeSafe-style
/v1/systemone endpoint. This is not Jev's model or architecture and is not affiliated with
TypeSafe AI.
Use the JevK5 runtime shown below to read option probabilities. Generic text-generation examples
on the Hub call generate() and do not perform JevK5's decision readout.
jevk5_config.json: temperature 1.22 and knockout_temperature 0.93v0.2 (see Use)knockout_temperature) from jevk5_config.json. Up to 16 options nothing changed: jevk5 0.2.2
and 0.3.0 give probabilities bit-identical to the benchmark run below on all 231 public
JevBench items. With 0.2.2, questions with more than 16 options fall back to v0.2's 0.77.The dev set is held-out rows only: teacher questions from three domains that training never saw (residential leases, public-sector permits, manufacturing QC), a hand-written hard set, and a hashed 5% of every public train split.
| v0.2 | v0.3 | |
|---|---|---|
| Index proxy (16 sources, see below) | 0.620 | 0.731 |
| Held-out teacher questions (362), accuracy | 0.801 | 0.834 |
| Hand-written hard set (64), accuracy | 0.766 | 0.766 |
| ECE on the teacher questions, calibrated | 0.034 | 0.035 |
Both columns are read the same way, through the runtime, on the same dev set. (v0.2's own card reports 0.804 on the teacher questions, from its training-time evaluation.)
The index proxy is our own estimate, not an index score. It is the chance-corrected skill, averaged over the 16 dev sources that are held-out train-split rows of Jev Decision Index benchmarks. Each source has 40 rows, so each one alone is noisy (about ±0.15). The largest gains are GSM8K (0.29 → 0.61), RAGTruth (0.15 → 0.50), SGD (0.68 → 0.93), When2Call (0.75 → 0.95) and iSarcasmEval (0.30 → 0.55). BANKING77 (0.91 → 0.88) and Amazon ESCI (0.50 → 0.47) slipped, and HoVer did not move (0.30). The hand-written hard set is level with v0.2 (49 of 64 for both).
On 2,500 hashed rows of the test split of avbiswas/bev-decision-150K (4,723 questions, every question type), v0.3 is level with v0.2: accuracy 0.663 against 0.665, ECE 0.036 against 0.036. By type: choice 0.697, yes/no 0.758, score 0.409. Leaving out the 34 questions whose document also appears in our training data changes accuracy by 0.002.
231 public items through JevBench's own runner (jevk5_direct adapter): 231/231 valid, 0
failures. The untrained row is the same base model and prompt without the LoRA or the temperature.
| Split | n | Untrained Qwen3.5-4B | JevK5 v0.2 | JevK5 v0.3 | v0.3 ECE (v0.2) |
|---|---|---|---|---|---|
| easy | 48 | 1.000 | 1.000 | 1.000 | 0.018 (0.038) |
| original (standard) | 72 | 0.986 | 0.958 | 0.944 | 0.057 (0.141) |
| hard (public half) | 111 | 0.613 | 0.739 | 0.784 | 0.054 (0.066) |
These use the runtime's knockout readout (groups of up to 16, then a final). The runs are 500 train-split items per dataset in the Decision Index's request shape, with every option offered. None of these items is in v0.3's training or dev data.
| Train split | Options | Passes | v0.2 accuracy / ECE | v0.3 accuracy / ECE | v0.3 macro-F1 |
|---|---|---|---|---|---|
| MASSIVE en-US (fitting set) | 60 | 5 | 0.754 / 0.038 | 0.738 / 0.045 | 0.724 |
| BANKING77 | 77 | 6 | 0.690 / 0.039 | 0.652 / 0.044 | 0.632 |
| CLINC150 with out-of-scope | 151 | 11 | 0.666 / 0.039 | 0.700 / 0.056 | 0.759 |
Teacher questions (17,408). A thinking model writes realistic documents with hard typed questions, then answers every question twice, independently. A question is kept only when both answers match the intended one. Option keys are rebuilt from the option text, so no key hints at the answer.
Public replay (30,052 items, train splits only).
| Dataset (train split) | License | Rows |
|---|---|---|
| GSM8K | MIT | 2,373 |
| WinoGrande (xl) | Apache-2.0 | 2,373 |
| HellaSwag | MIT | 2,373 |
When2Call (train_pref) | CC BY 4.0 | 1,978 |
| HoVer (claims + the cited Wikipedia introductions) | CC BY-SA 4.0 | 1,582 |
| iSarcasmEval (task A and task C formats) | MIT | 1,581 |
| ARC (Easy + Challenge) | CC BY-SA 4.0 | 1,186 |
| CommonsenseQA | MIT | 1,186 |
CLINC150 (plus) | CC BY 3.0 | 1,186 |
| Schema-Guided Dialogue (30 of 127 train files) | CC BY-SA 4.0 | 1,186 |
| Amazon ESCI (3 of 11 shards) | Apache-2.0 | 1,186 |
| RAGTruth | MIT | 1,186 |
| ToolACE | Apache-2.0 | 1,186 |
| WANLI | CC BY 4.0 | 1,186 |
| AQuA-RAT | Apache-2.0 | 791 |
| OpenBookQA | Apache-2.0 | 791 |
| Cosmos QA | CC BY 4.0 | 791 |
| SWAG | MIT | 791 |
| BoolQ | CC BY-SA 3.0 | 791 |
| MultiNLI | OANC / CC BY 3.0 / CC BY-SA 3.0 / MIT, by genre | 791 |
| BANKING77 | CC BY 4.0 | 791 |
| MASSIVE (en-US) | CC BY 4.0 | 791 |
| New Yorker caption contest (matching, fold 0) | CC BY 4.0 | 791 |
| QASC | CC BY 4.0 | 395 |
| RuleTaker | Apache-2.0 | 395 |
| Glaive function calling v2 | Apache-2.0 | 395 |
Two development variants are not released because of their data's terms:
auxiliary_train, most of them RACE reading
passages (RACE is for non-commercial research only). With the same readout it scored 0.740 on
the index proxy (this model 0.731), 0.820 on the teacher questions (0.834) and 0.781 on the
hand-written set (0.766).Training and calibration. Cross-entropy on the option-letter logits, SemIf's prompt format, 1 epoch, learning rate 3e-5, inputs up to 2,048 tokens. A question whose answer is a distribution trains against that distribution. One temperature is fitted on the held-out teacher questions: ECE 0.050 → 0.035.
Data rules.
Declared overlap with the Jev Decision Index. These are train splits of index benchmarks, deduplicated against their test and validation items: ARC, OpenBookQA, CommonsenseQA, GSM8K, WinoGrande, HellaSwag, BANKING77, CLINC150, SGD, Amazon ESCI, When2Call, iSarcasmEval, RAGTruth, HoVer and the New Yorker caption contest. Separately, 53 of v0.2's Qwen-written training questions share at least one 8-word sequence with ContractNLI (41) or SGD (12) test or dev text. They were reported by the scan and kept. v0.3 was not trained on any split of MMLU, MMLU-Pro, ANLI or NLI4CT.
from jevk5 import JevK5
model = JevK5("alibiserikbay/JevK5")
model.decide(
"I was billed twice for order #4411. Please refund the duplicate charge today.",
{"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "tech": "Bugs", "sales": "New purchases"}},
)
# {'type': 'choice', 'choice': 'billing', 'confidence': 0.996, 'probabilities': {...}, ...}
Or as a server that answers TypeSafe-style /v1/systemone requests:
jevk5-serve --model alibiserikbay/JevK5 --port 8090. Both read temperature and
knockout_temperature from this repo's jevk5_config.json.
For v0.2, pass a local copy of the tagged revision:
from huggingface_hub import snapshot_download
model = JevK5(snapshot_download("alibiserikbay/JevK5", revision="v0.2"))
Qwen3.5-4B and Qwen3.6-27B by the Qwen team (Apache-2.0). GPT-6 Luna by OpenAI. The one-pass readout and prompt come from SemIf by TheoLeeCJ (MIT). The public datasets in the table above belong to their authors, under the licenses listed. Evaluated with JevBench (github.com/fstandhartinger/jevbench, MIT) and bev-decision-150K. Not affiliated with TypeSafe AI or Jev.