Downloads · 30 days
82
100% of all-time downloads
oraculumai/Manchego
Manchego is a text classification model from oraculumai. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
<p align="center"<img src="manchego.jpg" alt="The Manchego mascot: a smiling wedge of Manchego cheese in a black beret, waving" width="560"</p
Downloads · 30 days
82
100% of all-time downloads
All-time downloads
82
Public
Parameters
4.7B
18.7 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors9.3 GB · 100%
How the weights are stored.
BF164.7B · 100%
From the Hugging Face model README
A 4B model for typed decisions. One forward pass, a probability for every option.
Give Manchego a state (text or JSON), a question and a closed set of options. It returns a probability for each one:
pick an option (choice), a yes/no condition (noul), or an ordered level (score). It does not generate text.
This is Manchego v3: Qwen3.5-4B with a LoRA adapter (merged here), trained from the untrained base under the SemIf
prompt on about 100M tokens of decision rows. Manchego v2.1, the previous release, stays available at tag v2.1.
What changed in v3
"contract": "semif", manchego-serve 0.2); v2.1's short prompt
is not v3's.| Suite (our runs, not official JevBench scores) | rows | untrained Qwen3.5-4B, v2.1's prompt | untrained Qwen3.5-4B, SemIf prompt | Manchego v2.1 | Manchego v3 | v3 minus v2.1 [95% interval] |
|---|---|---|---|---|---|---|
| JevBench, 231 public decisions (our run, not an official JevBench score) | 231 | 0.775 (0.493) | 0.801 (0.485) | 0.805 (0.423) | 0.857 (0.326) | +0.052 [+0.004, +0.102] |
| Public8, eight public real-text datasets | 1,600 | 0.686 (0.843) | 0.725 (0.847) | 0.732 (0.684) | 0.753 (0.609) | +0.021 [+0.002, +0.039] |
| 27 held-out Natural Instructions tasks (source-clean) | 1,561 | 0.678 (0.919) | 0.678 (1.056) | 0.701 (0.929) | 0.709 (0.941) | +0.007 [-0.016, +0.031] |
Accuracy with cross-entropy in brackets, temperature 1, one runtime for every model (see How the comparison was run). The JevBench rows on this card are our runs on its 231 public decisions, not official JevBench scores. JevBench's official score also reads sealed items, calibration, speed and cost, and only its maintainer runs it. v3 has no official JevBench result yet.
Sealed sets, read once. Two sets held back for this read, with no recorded Manchego training or scoring on either, were opened once, on 2026-09-30, after v3 was fixed; nothing was selected by that read. The unseen-task set's record also notes earlier exposure: before it was reserved, Jev (TypeSafe AI's hosted decision service) was run on training rows of its tasks, and researchers saw some of its task names.
| Sealed set, read once | rows | untrained Qwen3.5-4B, v2.1's prompt | untrained Qwen3.5-4B, SemIf prompt | Manchego v2.1 | Manchego v3 | v3 minus v2.1 [95% interval] |
|---|---|---|---|---|---|---|
| 24 unseen Natural Instructions tasks (13 source clusters) | 5,397 | 0.641 (0.918) | 0.674 (0.911) | 0.659 (0.849) | 0.695 (0.861) | +0.042 [-0.018, +0.107] |
| Fresh out-of-distribution draws of four of the project's own computed families | 1,193 | 0.445 (1.161) | 0.508 (1.091) | 0.795 (0.466) | 0.762 (0.578) | -0.033 [-0.060, -0.005] |
The model columns pool rows. The unseen-task contrast is the mean over the 24 tasks of the per-task difference, with whole source clusters resampled; v3 is also +0.060 [-0.020, +0.143] above the untrained base there, again not resolved. The family draws (policy decisions, lookup chains, temporal reasoning, answer adequacy; 400 instances resampled) are where v2.1 kept training on v2's families and v3 started over from the base: v3 is 0.033 below v2.1 (resolved) and +0.317 [+0.280, +0.353] above the base.
| JevBench v1.5.4, official (maintainer-run) | score (95% interval) | rank | Intelligence | Calibration | Speed | Cost |
|---|---|---|---|---|---|---|
Manchego v2.1 (the previous release, tag v2.1) | 68.8 (59.6 to 70.3) | 11 of 106 | 51.2 | 84.9 | 88.6 | 64.3 |
Measured by JevBench's maintainer on their own offline GPU (RTX6000, bf16, temperature 1.0, original option order) with
manchego-serve v0.1.1 and oraculumai/Manchego at the v2.1 commit 77403228; the headline score weighs the four axes
equally. This row is v2.1's, not v3's. v3 has no official result yet.
v3 is served by manchego-serve 0.2 (Apache-2.0). The server speaks
TypeSafe AI's System One wire contract (POST /v1/systemone), runs offline, and hashes the weights it loads and reports
whether they are the published ones. Pin this repository by its tag v3 (manchego-serve 0.2.0 pins the commit behind it).
git clone https://github.com/nschlaepfer/manchego-serve && cd manchego-serve && git checkout v0.2.0
# Docker, Linux + NVIDIA GPU: the build downloads the pinned weights once; the container runs offline
docker build -t manchego-serve:3-cuda --build-arg MODEL_REVISION=v3 .
docker run --rm --gpus all -p 127.0.0.1:8000:8000 manchego-serve:3-cuda \
--contract semif --temperature-map /models/manchego/temperature_map.json
# or pip (install the CUDA build of torch first on Linux + CUDA)
pip install .
manchego-serve-download --repo oraculumai/Manchego --revision v3 --out ./manchego-v3 # uses the network once
manchego-serve --model ./manchego-v3 --revision v3 --contract semif \
--temperature-map ./manchego-v3/temperature_map.json # offline from here on
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"model": "manchego-3",
"state": "Customer: the blender I bought last week smells of burning and stopped working. Order 5521.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this message?",
"criteria": {"returns": "refunds, exchanges and defective items",
"shipping": "delivery status and lost parcels", "billing": null}},
"defective": {"type": "noul", "instructions": "Does the customer report a defective product?"},
"urgency": {"type": "score", "instructions": "How urgent is this message?",
"criteria": ["routine", "soon", "immediately"]}}}'
answers.route.probabilities holds one probability per option, answers.defective.noul is P(yes) and
answers.urgency.score is the expected level. Every response names the prompt each question got
(manchego.contract_by_question), the temperatures used and the weights' hash. manchego_config.json in this
repository already selects semif; --contract semif states it. Use --temperature-map off in place of the map
for probabilities at temperature 1, the setting of every number on this card.
Without the server (Transformers). Not a pipeline("text-classification") model: read the option-letter logits.
contract_semif.py in this repository renders the prompt v3 was trained on (2 to 16 options; the server handles 17 to
255 with the state-first prompt of contract_v2.py).
# pip install "transformers==5.17.0" torch huggingface_hub (tested: transformers 5.17.0, torch 2.10.0)
import importlib.util, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO, REV = "oraculumai/Manchego", "v3"
spec = importlib.util.spec_from_file_location("contract_semif", hf_hub_download(REPO, "contract_semif.py", revision=REV))
semif = importlib.util.module_from_spec(spec); spec.loader.exec_module(semif)
tok = AutoTokenizer.from_pretrained(REPO, revision=REV)
model = AutoModelForCausalLM.from_pretrained(REPO, revision=REV, dtype=torch.bfloat16, device_map="auto").eval()
def decide(state, question, options, kind): # options: [(value, description or None)]; noul: [("true", None), ("false", None)]
r = semif.render(state, question, options, kind)
text = tok.apply_chat_template(r["messages"], tokenize=False, add_generation_prompt=True, enable_thinking=False)
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
logits = model(**ids).logits[0, -1].float()
z = logits[[tok.encode(L, add_special_tokens=False)[0] for L in r["letters"]]]
return dict(zip(r["keys"], torch.softmax(z, 0).tolist())) # divide z by the type's temperature to match the map
print(decide("WIN a free cruise! Reply YES now.", "Is this message spam?", [("true", None), ("false", None)], "noul"))
Apple silicon: Manchego-MLX-8bit and
Manchego-MLX-4bit.
| Question type | serving temperature (temperature_map.json) | --temperature-map off |
|---|---|---|
| choice | 1.5 | 1.0 |
| noul | 0.2 | 1.0 |
| score | 1.0 | 1.0 |
The probabilities are softmax(z / T) over the offered option codes, with T set by the question type. A temperature never changes the chosen option.
The noul temperature 0.2 sharpens yes/no probabilities on purpose. JevBench v1.5 counts a yes/no answer whose P(yes) lies between 0.20 and 0.80 as wrong, so the map pushes answers out of that band. Under the map, P(yes) is therefore not a calibrated probability wherever the model is unsure. The choice temperature 1.5 flattens choice probabilities; score is unchanged.
Every number on this card is at temperature 1. --temperature-map off serves exactly that.
How it was chosen (TMAP-V15). Per question type, the temperature from a grid (0.2 to 2.0) that maximises an estimate of JevBench v1.5's composite, computed with v1.5's published scoring rules, on v3's records of development rows it never trained on. No JevBench item chose it:
| TMAP-V15 fit | choice | noul | score | total |
|---|---|---|---|---|
| development rows | 10,485 | 16,050 | 1,647 | 28,182 |
The rule was registered before the fit, but after we had seen a noul-temperature sweep on those records and a report on JevBench's public items. After the fit, its effect on the public items was reported and changed nothing.
Which weights. The map is bound by hash to the three v3 builds: this repository's bf16 weights (weights_sha256 in
Files), where it was fitted, and the MLX 8-bit and 4-bit builds, where it is applied as-is. manchego-serve 0.2 ships it
as the default map for those hashes only; any other weights get temperature 1. It was fitted on records scored with the
adapter applied to the base in padded batches, not on these merged weights' one-prompt logits; the merge moves option
logits by at most 0.125 on its check rows.
v3's manchego_config.json turns on one piece of manchego-serve 0.2's CUDA fast path by default: the fast host path,
which is bit-for-bit the reference arithmetic with less host work (torch backend on CUDA only; --no-fast-host turns it
off). CUDA graphs stay off by default. On Linux, which is what the Docker recipe runs, padding each prompt to a graph
size made every prompt length slower:
| median latency, one decision per request (A10, Linux, Docker image, these weights, SemIf prompt) | reference | CUDA graphs + fast host |
|---|---|---|
| up to 256 prompt tokens | 54 ms | 71 ms |
| up to 512 | 134 ms | 143 ms |
| up to 1,024 | 190 ms | 275 ms |
| up to 2,048 | 357 ms | 543 ms |
On Windows (an RTX 5090) the same graphs were about five times faster than the reference (median 26 vs 124 ms), because
per-call overhead dominates there; --cuda-graphs turns them on if that is your platform. They pad each prompt to 256,
512, 1,024 or 2,048 tokens and replay a captured forward: on 83 invented test questions no chosen option changed, but
probabilities moved by up to 0.0385 (0.0626 across other bucket sets). Details: manchego-serve's docs/FAST_PATH.md. The
official Speed of 88.6 above is v2.1's, measured on the reference path.
contract_v2.py, which v3 never trained on. On the 104 rows of the held-out task set with
more than 16 options, v3 and v2.1 each got 78 right.manchego_config.json.| Base | Qwen/Qwen3.5-4B @ 851bf6e8 (hybrid gated-delta-net + attention), untrained |
| Method | LoRA, rank 16, alpha 32, on q/k/v/o of the 8 attention layers and the five in/out projections of the 24 gated-delta-net layers (152 modules, 14,376,960 trainable parameters), merged into the base for release |
| Objective | cross-entropy on the offered option letters' logits at the last prompt position, with exact soft targets where a row declares a distribution; no loss on any text token |
| Prompt | SemIf's direct-options prompt on every row (2 to 16 options); choice options shuffled at every presentation, noul true then false |
| Run | 255,991 rows in one pass, 8,771 updates (token-budget batches of 12,288 padded tokens, at most 64 rows), 99,362,993 prompt tokens; lr 3e-5, 85 warm-up updates, cosine to 0, AdamW, no weight decay, gradient clip 1; one NVIDIA GH200, 2026-09-28, 00:56 to 04:48 UTC; every planned row consumed |
| Checkpoint | the final update; nothing was selected inside the run |
| Training data (bucket) | share of the estimated prompt tokens | rows | labels |
|---|---|---|---|
| Computed decision families of v2 (policies, lookups, temporal, adequacy, routing, skills, arena and Doom states, score levels), regenerated | 27.3% | 48,914 | computed by our generators |
| Hard computed families (long policy documents, multi-hop lookups, abstention, temporal strata, exact probabilities, paraphrase invariance, policy compliance) | 22.7% | 19,214 | computed |
| Answer judging (judging worked answers; judge-error repair) | 8.0% | 22,547 | computed |
| Human-annotated text: WANLI, MultiNLI, Bitext; 101 Natural Instructions tasks; GSM8K, MathDial, Spider, ToolACE, When2Call, NVD/CVE/CWE, ProsocialDialog, deepset prompt-injections | 20.0% | 83,701 | the datasets' original annotation (ToolACE and When2Call: their authors' generated labels) |
| Self-distilled replay (a partition of the same public corpora, and fresh answer-judging instances) | 15.0% | 61,739 | the untrained Qwen3.5-4B's own probabilities over the options |
| Domain generators (auth-log triage, code output, exact draws, genetic crosses, limitation deadlines, loan amortisation, math-answer grading, plan cost sharing) | 5.0% | 11,087 | computed |
| Calibration rows | 2.0% | 8,789 | computed |
| Total | 100% | 255,991 |
By target, 178,334 rows train toward one gold option, 15,918 toward an exact distribution and 61,739 toward the base's own probabilities. By type: 158,909 choice, 84,115 noul and 12,967 score rows.
No row comes from a benchmark on this card. Every source went through an 8-word overlap gate against the evaluation sets (JevBench's public items, Public8, the held-out and sealed task sets) before the training file was built, and the rows that matched were dropped. Public8's eight datasets and the held-out tasks' upstream datasets appear nowhere in the declared training data (checked by name); what the base model saw in pretraining is unknown. GSM8K and When2Call are also Decision Index benchmarks: v3 trained on their train splits only, and their test splits were kept out of training. "Computed" means exact under the task definition as implemented.
| Calibration error (ECE, temperature 1) | untrained Qwen3.5-4B, v2.1's prompt | untrained Qwen3.5-4B, SemIf prompt | Manchego v2.1 | Manchego v3 | v3 mean confidence | v3 accuracy |
|---|---|---|---|---|---|---|
| JevBench, 231 public decisions (our run, not an official JevBench score) | 0.062 | 0.059 | 0.072 | 0.055 | 0.854 | 0.857 |
| Public8 | 0.122 | 0.122 | 0.053 | 0.055 | 0.809 | 0.753 |
| 27 held-out Natural Instructions tasks | 0.107 | 0.119 | 0.083 | 0.129 | 0.829 | 0.709 |
| 24 unseen Natural Instructions tasks (sealed) | 0.151 | 0.148 | 0.098 | 0.141 | 0.836 | 0.695 |
| Fresh out-of-distribution draws of the project's families (sealed) | 0.141 | 0.075 | 0.051 | 0.050 | 0.800 | 0.762 |
On familiar ground v3 is as well calibrated as v2.1. On unfamiliar task definitions it is overconfident: on the 27 held-out tasks its mean confidence is 0.829 against an accuracy of 0.709, and its calibration error (0.129) is worse than v2.1's (0.083) and the base's. v2.1's second stage had repaired exactly this in v2; v3, trained from the base, has it again. The serving map's choice temperature of 1.5 flattens choice probabilities; its effect on these sets was not measured.
--temperature-map off when
you need P(yes) as a probability.Every model in the tables above ran in one runtime on one RTX 5090: bf16, the adapters applied to the base (not merged),
padded batches of up to 8,192 tokens, identical rows, temperature 1. v3 ran under the SemIf prompt (the state-first
prompt beyond 16 options, as manchego-serve 0.2 serves it); v2.1 and the untrained base ran under v2.1's served policy
(the short prompt up to 26 options, state-first beyond), and the base also under SemIf. v2.1 and the untrained base
reproduced their archived records on the three public suites (the runtime check passed). Intervals are paired 95%
bootstraps: JevBench resamples its scenario groups, Public8 its items within each dataset, the held-out set whole tasks.
The merged weights in this repository reproduce base plus adapter within the merge check (merge_record.json: at most
0.125 in any option logit on 12 rows, no changed decision); a bf16 merge is not bit-identical. Aggregates:
eval/.
adapter_model.safetensors
sha256 c43e2688… (full hash in merge_record.json). Training file sha256 240b5cbc…, 255,991 rows, checked row by
row on the training machine by a launch gate: no reserved evaluation task; no target from an external teacher model
or a hosted decision service (the replay's targets are the base model's own); every row's permission recorded.NOTICE carries every attribution those permissions require.eval/de_minimis_probe.json): Bitext, MultiNLI fiction and Spider pass every
rule. MathDial fails two: under the decision prompt, 0.075 of trained dialogues are continued verbatim for 8 tokens or
more against the base's 0.045 (the rule allows 0.02 more), and 6 trained dialogues are continued for 12 tokens or more
by the adapter alone (the rule allows none; as raw text, 2). On 200 MathDial dialogues it never trained on, the same
measures read 0.155 against 0.085, and 14.model.safetensors-* in this repository are the base checkpoint with every language-model tensor replaced by
its merged value; the vision tower, the multi-token-prediction head, the config and the tokenizer are the base
model's. Manchego was trained and evaluated on text only.This repository's main holds Manchego v3 (2026-09-30), tagged v3. Manchego v2.1 (2026-09-21) stays available,
unchanged, at tag v2.1 (revision="v2.1"): its weights, card, DETAILS.md and evidence. v2.1 needs its own prompts
(manchego-serve's contract auto); v3 needs SemIf. Earlier versions are not published.
| Format | repository | size | weights_sha256 (as manchego-serve reports it) |
|---|---|---|---|
| bf16 (Transformers) | oraculumai/Manchego | 9.3 GB | 2ee838433bfe278a226dc644667ad4a99ece82cc47325c7645a7dae723c1863b |
| MLX 8-bit | oraculumai/Manchego-MLX-8bit | 4.5 GB | 358b025b04001e50a065f8c87929175211264bd6182af74b67bd6caa2f639657 |
| MLX 4-bit | oraculumai/Manchego-MLX-4bit | 2.4 GB | e1bc5538b8dced2a857b4980dba045c2ca01db1c369aa416fc19f0f5c593e782 |
Each conversion is its own numeric series; the MLX cards carry their own measurements. No GGUF build is published.
Also here: LICENSE (Apache-2.0), NOTICE (every attribution and declaration), NATURAL_TASKS_ATTRIBUTION.md (the
101 Natural Instructions tasks and their 42 upstream sources), manchego_config.json (the decision contract; manchego-serve
reads its contract), temperature_map.json (the serving map), contract_semif.py (the SemIf prompt, 2 to 16 options),
contract_v2.py (the state-first prompt, 17 to 255 options), merge_record.json and eval/ (the aggregates behind this
card).
Base model: Qwen3.5-4B (Apache-2.0, Alibaba Cloud). Training corpora, each used under its own licence: WANLI (Liu,
Swayamdipta, Smith, Choi, 2022; CC BY 4.0); MultiNLI (Williams, Nangia, Bowman, 2018; mostly under the OANC licence;
in the fiction genre one work is CC BY-SA 3.0, two are CC BY 3.0, the rest US public domain); Bitext customer support
dataset (Bitext Innovations; CDLA-Sharing-1.0); GSM8K (Cobbe et al., 2021, OpenAI; MIT); MathDial (Macina et
al., 2023, ETH Zurich; CC BY-SA 4.0); Spider (Yu et al., 2018, Yale LILY; CC BY-SA 4.0); ToolACE (Liu et al.,
2024, Team-ACE; Apache-2.0); When2Call (Ross et al., 2025, NVIDIA; CC BY 4.0); NVD/CVE/CWE (NIST's National
Vulnerability Database; CVE and CWE content copyright The MITRE Corporation, used under MITRE's terms); ProsocialDialog
(Kim et al., 2022, Allen Institute for AI; CC BY 4.0); deepset prompt-injections (deepset; Apache-2.0);
Super-NaturalInstructions (Wang, Mishra, et al., 2022; Apache-2.0 collection; each task's instances under its upstream
dataset's licence, all listed in NATURAL_TASKS_ATTRIBUTION.md, where the one Gigaword-based task is noted). No row of
any corpus is redistributed here. Doom states were recorded from ViZDoom (MIT) scenarios with Freedoom assets (BSD-3).
The prompt format and system message are SemIf's direct-options prompt (SemIf, formerly OpenJev, by TheoLeeCJ; MIT).
Evaluation sets: JevBench, the eight-dataset suite of logan-markewich/jeff (MIT) we call Public8, and
Super-NaturalInstructions, each dataset under its own licence. Full notices: NOTICE. Illustration: the Manchego mascot,
an AI-generated image (ChatGPT image generation) supplied by the project's author. The interface follows TypeSafe AI's
System One contract; Manchego is an independent project, not affiliated with or endorsed by TypeSafe AI.