Downloads · 30 days
549
100% of all-time downloads
kirp/jpt-0.8b
jpt-0.8b is a image-text-to-text model from kirp. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
[](❯❯-license) [](❯-jevbench-v140) [](❯-decision-index-021) [](https://github.com/tic-top/llm2jev)
Downloads · 30 days
549
100% of all-time downloads
All-time downloads
549
Public
Parameters
853M
1.7 GB on disk
Likes
3
Trending 1
Click a slice to open those files.
.safetensors1.7 GB · 98%
How the weights are stored.
BF16853M · 100%
From the Hugging Face model README
JPT-0.8B · JPT-4B · JPT-9B · JPT-35B-A3B · llm2jev
JPT-0.8B is a fast, open decision model: give it a situation and typed questions, get a calibrated probability for every option from one forward pass. No generated explanation, no reasoning tokens — latency is one prefill.
The small sibling of JPT-4B: the same recipe and the same data, small enough for CPU or edge serving.
It implements the typed-decision interface introduced by Jev from TypeSafe AI [1]: a caller sends a state plus questions, and each question is one of three types. JPT is an independent model, not derived from Jev and not trained on Jev outputs; it is an open alternative behind the same interface.
| Question type | What it answers | Options |
|---|---|---|
choice | pick one | 2–255 labels |
score | a level on an ordered scale | the scale's levels |
noul | yes / no | true, false |
Built on Qwen/Qwen3.5-0.8B: a LoRA fine-tune merged into full weights. The vision tower is unchanged.
Probabilities use one temperature T = 1.140, fit once on a held-out split — never per benchmark.
JevBench [2] scores general typed decisions; JPT-0.8B reaches 0.736 public accuracy, above every sub-2B system in the v1.4 results and 0.113 above its base model. v1.4.0 has 231 public items and 308 sealed items only the maintainer can run, so JPT-0.8B has no official v1.4 score yet.

| System | Params | Public accuracy (231) |
|---|---|---|
| JPT-35B-A3B (ours) | 35B-A3B | 0.892 |
| JPT-4B (ours) | 4B | 0.879 |
| Jev 1.13.0 (TypeSafe AI, API) | closed | 0.866 |
| JPT-9B (ours) | 9B | 0.853 |
| JPT-0.8B | 0.8B | 0.736 |
| decider-2b | 2B | 0.710 |
| kev 0.6B | 0.6B | 0.667 |
| Open-Jev 2B | 2B | 0.645 |
| Qwen3.5-0.8B, same prompt, zero-shot (our run) | 0.8B | 0.623 |
| SimpleJev (Qwen3.5-0.8B) | 0.8B | 0.545 |
| kev 0.5B | 0.5B | 0.494 |
Source: other rows from results/v1.4/jevbench-v1.4-results.json at jevbench commit 2fa63fa (2026-09-23); ours
scored with the benchmark's own CLI.
Decision Index [3] 0.2.1 (2026-09-27) is the broadest test: the full frozen suite, 38 scored benchmarks in five areas, chance-corrected — JPT-0.8B scores 19.22, the best 0.8B on the board. Run through llm2jev over SGLang and submitted as apolinario/decision-index#7; 0.2.1 rescores the same run (17.07 under 0.2).

| Model | Params | Decision Index 0.2.1 |
|---|---|---|
| Jev 1.13.0 (TypeSafe AI, API) | closed | 57.91 |
| JPT-35B-A3B (ours, not on the board yet) | 35B-A3B | 52.89 |
| JPT-0.8B | 0.8B | 19.22 |
| Decision 1.0 Eos | 0.8B | 18.41 |
| Kev 0.8B | 0.8B | 14.60 |
| MoJev | 0.8B | 11.69 |
Source: live board data/index-v0.2.1.json (generated 2026-09-27 16:59 UTC); full run and scores.json in
kirp/decision-index-results-jpt-0.8b
(gated: it carries the suite's GPQA/HLE item text).
By area, against Jev 1.13.0 on the same items and scorer: JPT-0.8B is behind Jev in all five areas and ahead of it on 1 of the 38 index benchmarks.

| Area | Jev 1.13.0 | JPT-0.8B |
|---|---|---|
| Knowledge | 51.4 | 9.9 |
| Language | 62.0 | 23.0 |
| Retrieval | 55.4 | 16.4 |
| Tools | 75.1 | 36.2 |
| Arts | 37.7 | 8.0 |
| Area | Benchmark | Jev 1.13.0 | JPT-0.8B |
|---|---|---|---|
| Arts | BPoMP | 81.8 | 0.0 |
| Arts | ForecastBench | 30.6 | 5.8 |
| Arts | Habermas Machine | 21.5 | 12.1 |
| Arts | Humicroedit | 23.7 | 4.4 |
| Arts | New Yorker | 62.6 | 25.2 |
| Arts | POP909-CL | 15.9 | 1.7 |
| Arts | cfcolor | 28.8 | 7.2 |
| Games | ChessBench | 9.8 | 0.3 |
| Knowledge | BBH | 89.7 | 21.2 |
| Knowledge | CLadder | 45.3 | 8.8 |
| Knowledge | CRUXEval | 57.1 | 0.0 |
| Knowledge | GPQA Diamond | 71.4 | 6.8 |
| Knowledge | GSM8K | 75.6 | 20.7 |
| Knowledge | HLE | 4.7 | 0.9 |
| Knowledge | MMLU-Pro | 80.5 | 19.0 |
| Knowledge | MuSR | 46.1 | 16.7 |
| Knowledge | SATA-Bench | 25.4 | 2.7 |
| Language | ACOS | 27.3 | 4.4 |
| Language | ANLI | 62.2 | 8.0 |
| Language | ContractNLI | 59.1 | 60.6 |
| Language | FinEntity | 80.8 | 67.7 |
| Language | HellaSwag | 92.7 | 31.5 |
| Language | NLI4CT | 69.0 | 27.7 |
| Language | RAGTruth | 51.3 | 7.4 |
| Language | VAST | 46.9 | 16.0 |
| Language | WinoGrande | 83.9 | 6.4 |
| Language | iSarcasmEval | 36.3 | 5.5 |
| Retrieval | Amazon ESCI | 43.8 | 12.3 |
| Retrieval | BANKING77 | 79.5 | 34.6 |
| Retrieval | BRIGHT | 40.6 | 23.9 |
| Retrieval | CLINC150+OOS | 89.2 | 9.9 |
| Retrieval | HoVer | 45.7 | 9.7 |
| Retrieval | PhishNChips phishing decisions | 25.1 | 4.1 |
| Tools | API-Bank | 88.0 | 28.8 |
| Tools | BFCL | 94.3 | 68.6 |
| Tools | Home appliance simulator | 52.3 | 0.0 |
| Tools | ToolRet | 59.9 | 45.6 |
| Tools | When2Call | 74.6 | 33.3 |
Chance-corrected skill × 100 (0 = random, 100 = perfect). Jev's numbers are its official entry on the live board
(jev-1.13.0); ours are from the same kit and suite.
The same evals as the larger JPTs, at T = 1.140. Same-prompt base-model rows have not been run for these at 0.8B.
| Benchmark (version, n) | What it tests | JPT-0.8B | Jev 1.13.0 |
|---|---|---|---|
| JevBench v1.4.0 public hard tier [2] (111) | hardest general decisions | 0.577 (ECE 0.156) | — |
| Typed decisions test (ours, 2,000) | in-distribution typed decisions | 0.768 (ECE 0.162) | — |
| ANLI r1 / r3 [4] (dev) | adversarial NLI | 0.437 / 0.483 | — |
| Banking77 [5] / MASSIVE 1.1 [6] (en / de / zh) | intent classification | 0.650 / 0.797 / 0.697 / 0.760 | — |
| EnvBench v0.1 (ours) public / held-out [10] (skill 0–100) | sequential decisions in game envs | 37.8 / 35.0 | — |
| ScreenSpot-v2 [7] SoM / Screen2Words [8] / ERQA [9] (images, zero-shot) | GUI grounding, screen summary, embodied reasoning | 0.787 / 0.740 / 0.302 | — |
Jev 1.13.0 has no official score on these splits (our own test/dev cuts, EnvBench, and the image sets), so its column is "—"; its official scores on the Decision Index versions of ANLI and BANKING77 are in the per-benchmark table above. Running the Jev API on these splits would fill them.
Banking77, MASSIVE and typed rows are in-distribution (train splits in the mix, test items not).
Two pieces: an engine that holds the weights, and llm2jev (>= 0.6.1) in front of it, reading option probabilities off the engine.
python -m sglang.launch_server --model-path kirp/jpt-0.8b --port 30000 \
--context-length 32768 --mamba-scheduler-strategy extra_buffer & # Qwen3.5's DeltaNet layers need this flag
llm2jev --model kirp/jpt-0.8b --backend sglang --url http://127.0.0.1:30000 --port 8080 --temperature 1.140
Tested with SGLang 0.5.9; install cuDNN 9.15+ over its pinned 9.10:
pip install "sglang==0.5.9" && pip install "nvidia-cudnn-cu12>=9.15".
vllm serve kirp/jpt-0.8b --max-logprobs 256 --return-tokens-as-token-ids --enable-scale-out --port 8000
llm2jev --model kirp/jpt-0.8b --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.140
The three vLLM flags are required: without them every request is a bare HTTP 400.
pip install "llm2jev[hf,vision]"
llm2jev --model kirp/jpt-0.8b --backend hf --port 8080 --temperature 1.140
Serializes requests; fine on CPU at this size, for traffic use SGLang or vLLM.
import requests
r = requests.post("http://127.0.0.1:8080/v1/systemone", json={
"state": "Refund policy: full refund within 30 days of purchase; 50% until day 60; none after.\n"
"Order 1182 was bought on 3 March and returned on 20 April.",
"questions": {
"refund": {"type": "choice", "instructions": "What refund does order 1182 get?",
"criteria": {"full": "Full refund", "half": "50% refund", "none": "No refund"}},
"late": {"type": "noul", "instructions": "Was the return made after day 30?",
"criteria": {"true": "Yes", "false": "No"}}}})
print(r.json()["answers"]) # each answer has the per-option probabilities
state = [{"role": "user", "content": [
{"type": "image", "image": "https://example.com/screen.png"},
{"type": "text", "text": "Task: open the settings page. Numbered boxes mark clickable elements."}]}]
questions = {"click": {"type": "choice", "instructions": "Which box should be clicked?",
"criteria": {"1": None, "2": None, "3": None, "4": None, "5": None}}}
The JPT-4B recipe and data (mix_train_env_v11, 49,221 typed questions) on a larger base, one epoch, merged into full weights.
| Part | What it is |
|---|---|
| Method | LoRA r=16 on every attention, DeltaNet and MLP projection of the language model, lr 5e-5; vision tower untouched |
| Loss | multi-class Brier over the option labels, on llm2jev's chat prompt with thinking disabled |
| Batch | 2 GPUs × 5 × 4 gradient-accumulation steps = 40 questions per step |
| Data | 49,221 questions in 32,835 records; one epoch over two option-shuffled copies — sources on the JPT-4B card |
| Held out | no item from JevBench, EnvBench held-out seeds, the Decision Index frozen suite or our typed test split |
CC BY-NC 4.0. The weights derive from Qwen3.5-0.8B (Apache-2.0), but some training datasets allow only non-commercial or research use, so the model is released for non-commercial use.
<small>JPT-0.8B is an independent open model that implements a typed-decision interface (noul, choice and score questions answered with probabilities). It is not affiliated with, endorsed by or derived from TypeSafe AI or its Jev model, and it was not trained on Jev outputs.</small>