Downloads · 30 days
0
Fr0zencr4nE/jev-spatial
jev-spatial is a machine learning model from Fr0zencr4nE. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Sep 24, 2026
Repo size
8.9 GB
Likes
1
Public
Click a slice to open those files.
.safetensors8.9 GB · 100%
From the Hugging Face model README
Fast spatial intelligence through finite-choice decisions
English · 简体中文
</div>Jev-Spatial is a System One spatial-intelligence model built on Molmo2-ER. Inspired by Jev, it gives spatial tasks one shared format: images + question + candidate options → one decision. All tasks go through one unified head, and the model never generates text.
Every task becomes a choice among a fixed set of options. The tasks fall into three families:
Every output comes from a fixed set of options, so there is nothing to parse and no malformed output. Compared with the autoregressive (AR) baseline, i.e. our reproduction of native Molmo2-ER, Jev-Spatial:
[!NOTE] Jev-Spatial is an independent research project. It follows the Jev idea of answering typed questions with choices instead of generated text. It is not affiliated with, endorsed by, or derived from TypeSafe AI or its Jev model, and it was not trained on Jev outputs.
Many everyday spatial questions are closer to System 1 perception than to deliberate reasoning:
A VLM already has the visual features and the image-text alignment needed to answer them. Making it describe its perception token by token, then parsing that text back into an answer, adds latency and new ways to fail.
Jev showed that many decisions can be read directly as a choice among typed options. Jev-Spatial applies the same idea to spatial perception. We hypothesize that restricting the output to a finite set of states also makes learning easier, because the model only has to rank a few options instead of producing a precise string. This may explain part of the pointing gains.
flowchart LR
A["image(s) + question<br/>+ candidate options"] --> B["Molmo2-ER backbone<br/>(LoRA-merged)"]
B --> C["unified head<br/>LayerNorm → Linear<br/>(invalid options masked)"]
C --> D{task}
D -->|classification| E["option ID"]
D -->|numeric regression| F["range → sub-range → meters"]
D -->|pointing| G["3×3 → 3×3 → 3×3 → (x, y)"]
F -. next round .-> B
G -. crop refill .-> B
All three task families end the same way: the unified head picks one option from a fixed set. They differ in what the options are, how many rounds it takes, and whether the image changes between rounds.
| Task family | answer_space.kind | Jev analogue | Options per round | Rounds | Image changes between rounds? | Output |
|---|---|---|---|---|---|---|
| Classification | choice | Yes/no · Choice | The 2–N options in the request | 1 | — | Option ID |
| Numeric regression | scalar | Score (ordered levels) | Value ranges, then sub-ranges | 2 | No | Length in meters |
| Pointing | point | (new) | 9 cells of a 3×3 grid | 3 | Yes (crop refill) | Normalized (x, y) in $[0, 1]$ |
Classification: one decision. This covers spatial relations, directions and yes/no questions. The request supplies the candidate options. They are shuffled so the model can't learn to prefer a position, and the model picks one in a single pass. This is Jev's native setting and needs no adaptation.
Numeric regression: pick a range, then narrow it down. A continuous value is split into ordered ranges. Round 1 picks a coarse range, with separate classes for exactly zero and above the maximum. Round 2 picks a sub-range inside it, and the prediction is decoded from that sub-range. The image and question stay the same across both rounds; only the options change. Accuracy therefore depends on how the ranges are designed (see Limitations).
Pointing: a special case. Pointing differs from the other two in two ways: its answer is a location in the image, and it is the only task where the visual input changes between rounds.
This zooming is what makes pointing work (see the ablation below). It is also where most of the extra compute goes.
max_choices options, and scores for options that don't exist in a request are masked out.--preserve-option-order to turn this off.We compared three ways to use the image across the three pointing rounds:
single_image: reuse the original image in every round, add a description of the region selected so far, and reuse its cache.roi_mask: reuse the original image's cache, but block the new tokens from attending directly to image tokens outside the selected region.crop_refill: crop the selected region, encode it again, and append it to the existing context.These are early checkpoints, not the released one. Each was trained for 300 steps and evaluated with 2-crop images and three 3×3 rounds. RefSpatial scores are region-hit rates on 200 questions.
| Variant | SAT real ↑ | VST MAE (m) ↓ | RefSpatial ↑ | Location ↑ | Placement ↑ |
|---|---|---|---|---|---|
single_image | 78.7 | 0.599 | 16.0 | 16.0 | 16.0 |
roi_mask | 77.7 | 0.603 | 18.5 | 21.0 | 16.0 |
crop_refill | 77.7 | 0.599 | 35.0 | 43.0 | 27.0 |
Summary: the choice of method barely affects classification or numeric regression. For pointing, crop_refill improves the region-hit rate by +19.0 points over single_image, while roi_mask improves it by only +2.5. The released model therefore uses crop_refill. All three variants are in runtime.py (point_variant).
<sub>Record: artifacts/benchmarks/fast-v1-20260923T201259Z/comparison.json</sub>
Molmo2-ER splits each image into local tiles and adds one global thumbnail, so it sees both fine detail and the overall layout. 24-crop means up to 24 local tiles plus the thumbnail. The actual number depends on image size and aspect ratio. Pointing crops go through the same preprocessing, so each extra round adds compute.
</details>| Data | ~72K QA pairs: SAT ~25K · VST-P ~22K · RefSpatial ~25K |
| Trainable parameters | LoRA on the language model + the unified head |
| Frozen | Vision encoder and the projector that connects it to the language model |
| Hardware | 8 × A800 |
| Release | LoRA merged into the backbone (no PEFT needed at inference) |
All image benchmarks use 24-crop and were run locally. Scores are percentages (↑ is better). VST reports mean absolute error in meters on 300 internal dev samples (↓ is better).
| Benchmark | Molmo2-ER (reproduced) | Naive three-head | Jev-Spatial | Δ vs. Molmo2-ER (reproduced) |
|---|---|---|---|---|
| SAT real ↑ | 79.3 | 77.7 | 75.3 | −4.0 |
| CV-Bench ↑ | 87.3 | 87.0 | 86.3 | −1.0 |
| RefSpatial-Bench ↑ | 52.5 | 9.0 | 54.5 | +2.0 |
| Where2Place ↑ | 57.0 | 26.0 | 64.0 | +7.0 |
| RoboSpatial-Pointing † ↑ | 29.5 | 4.1 | 59.8 | +30.3 |
| RoboSpatial-VQA † ↑ | 58.0 | 58.3 | 64.2 | +6.2 |
| VST dev MAE (m) ↓ | 0.520 | 0.428 | 0.472 | −9% error |
[!WARNING] † The RoboSpatial numbers are not yet verified. Our native AR VQA score (58.0) is well below the published Molmo2-ER result (73.4). Treat these rows as provisional until the evaluation is fixed.
Jev-Spatial is fastest when one image gets many independent questions, which is common for robots and agents. We encode the image once and share its cached context, run all questions in parallel, and batch together the pointing crops needed in the same round. Pointing always uses the full three 3×3 rounds.
Setup: 20 images and 160 questions (8 per image) from RoboSpatial, covering spatial relations, whether an object can be placed somewhere, and pointing to free space. Single A800, each configuration repeated 3 times. The table reports the mean time until all 8 questions about one image are answered.
| Method | Inference mode | 2-crop, ms ↓ | 24-crop, ms ↓ | Pointing hit rate, 24-crop ↑ |
|---|---|---|---|---|
| Molmo2-ER (reproduced) | one question at a time | 2848.0 | 5572.6 | 20.0% |
| Molmo2-ER (reproduced) | shared image, parallel | 717.9 | 1123.0 | 21.8% |
| Naive three-head | shared image, parallel | 148.3 | 527.1 | 7.3% |
| Jev-Spatial | one question at a time | 1397.6 | 4604.2 | 56.4% |
| Jev-Spatial | shared image, parallel | 675.7 | 1294.0 | 58.2% |
Speedup from sharing the image, for Jev-Spatial:
| Questions per image | 1 | 2 | 4 | 8 |
|---|---|---|---|---|
| 2-crop | 18.6% slower | 1.18× | 1.58× | 2.07× |
| 24-crop | 5.6% slower | 1.53× | 2.31× | 3.56× |
Summary
<sub>These speedups combine all three optimizations: sharing the image, running questions in parallel, and batching crops. In an early two-image test, crop batching alone saved only ~5.5% (24-crop, 8 questions per image: 1444.1 → 1365.0 ms), which is too small a test to count as a formal ablation. In BF16, parallel and one-by-one runs do not produce bit-identical outputs, so each quality number comes from that mode's own outputs. Record: artifacts/benchmarks/scene-latency-20260924/comparison.json</sub>
Milliseconds per request, single A800 after warmup. 20 fixed samples per benchmark, each run 3 times; we take the median per sample and average. Timing covers image loading, preprocessing, inference and output parsing. Image tasks use 24-crop; VST uses 2-crop.
| Benchmark | Molmo2-ER (reproduced) | Naive three-head | Jev-Spatial |
|---|---|---|---|
| SAT real | 504.3 | 461.2 | 484.3 |
| CV-Bench | 202.8 | 161.8 | 158.2 |
| RefSpatial-Bench | 741.4 | 127.9 | 306.1 |
| Where2Place | 1083.2 | 121.8 | 299.9 |
| RoboSpatial-Pointing | 986.8 | 464.8 | 772.9 |
| RoboSpatial-VQA | 542.5 | 463.2 | 471.1 |
| VST numeric dev | 300.1 | 111.8 | 187.4 |
| Mean (benchmarks weighted equally) | 623.0 | 273.2 | 382.9 |
With one question per request, classification speed is close to native AR. The biggest savings are on pointing benchmarks, where AR has to generate coordinate text: up to 3.6× faster on Where2Place. Reproduce with scripts/benchmark_latency.py.
Requires Python ≥ 3.10 and a CUDA GPU.
git clone https://github.com/Fr0zenCrane/jev-spatial
cd jev-spatial
pip install -e '.[inference]'
hf download Fr0zencr4nE/jev-spatial --local-dir models/jev-spatial
jev-spatial --model models/jev-spatial --input examples/requests.jsonl
--input takes a single .json request or a .jsonl file with one request per line. Image paths are resolved relative to the request file.
| Flag | Description |
|---|---|
--output PATH | Write results to a file instead of stdout |
--device | Default cuda:0 |
--max-crops N | Override the Molmo2-ER crop limit (the release defaults to 2; the image benchmarks use 24) |
--max-sequence-length N | Override the total token budget for all rounds |
--preserve-option-order | Don't shuffle classification options |
--seed N | Override the per-sample shuffle seed |
from spatial_jev.inference import JevSpatial
model = JevSpatial.from_pretrained("models/jev-spatial", device="cuda:0")
# Classification: spatial relations, directions, yes/no
model.classify("examples/scene.png",
"Where is the red square relative to the blue circle?",
["left", "right"])
# Numeric regression: nonnegative length in meters
model.measure("examples/scene.png", "How tall is the chair?", quantity="height")
# Pointing: one normalized (x, y) point in a single image
model.point("examples/scene.png", "Point to the blue circle.")
Image paths in the Python API are resolved relative to the current working directory.
// classification: 2..max_choices options, each with a unique, nonempty id and text
{"media": [{"kind": "image", "uri": "scene.png"}],
"question": "Where is the red square relative to the blue circle?",
"answer_space": {"kind": "choice",
"options": [{"id": "left", "text": "left"},
{"id": "right", "text": "right"}]}}
// numeric regression: this checkpoint estimates nonnegative lengths in meters
{"media": [{"kind": "image", "uri": "scene.png"}],
"question": "How tall is the chair?",
"answer_space": {"kind": "scalar", "quantity": "height", "unit": "m"}}
// pointing: exactly one image and one point
{"media": [{"kind": "image", "uri": "scene.png"}],
"question": "Point to the blue circle.",
"answer_space": {"kind": "point", "coordinate_system": "normalized_xy", "num_points": 1}}
| Field | Meaning |
|---|---|
prediction | Option ID, a value in meters, or an (x, y) point |
path | Index chosen at each round |
logits | Classifier scores for every option at each round |
mapping | How the shuffled options map back to the original ones (classification only) |
input_tokens | Total input tokens across all rounds |
src/spatial_jev/
├── inference.py # JevSpatial API + `jev-spatial` CLI
├── runtime.py # multi-round inference, unified head, point variants
├── hierarchy.py # scalar ranges and 3×3 grid encoding/decoding
├── schema.py # request checks and prompt building
├── unified.py # training model for the unified head
└── molmo2/ # bundled Molmo2 model and processor code (no remote code)
scripts/ # data prep, training, evaluation, latency, export
configs/ # pilot_v0 (three-head), unified_v1, mixed_v2
data/manifests/ # dataset and benchmark source lists
tests/
Special thanks to Molmo2-ER, which provides the spatial understanding, and to Jev, which inspired the decision-based approach.
We also thank Molmo2, MolmoAct2, Qwen, SigLIP 2, and the open Jev-style community projects jev-visual, Jev-Omni, Qwen-2.5-1B-RLCD, OpenJev, OmniJev, OpenJev-Vision, and SemIf.
Data and benchmarks: SAT, VST, RefSpatial / RoboRefer, CV-Bench, RoboPoint / Where2Place, RoboSpatial, VSI-Bench, and the original scene datasets. Tooling: PyTorch, Transformers, PEFT, Safetensors.
@misc{jevspatial2026,
title = {Jev-Spatial: Fast Spatial Intelligence through Finite-Choice Decisions},
author = {Fr0zenCrane},
year = {2026},
howpublished = {\url{https://github.com/Fr0zenCrane/jev-spatial}}
}
Code and weights: Apache-2.0. See NOTICE for third-party attributions. Datasets keep their own licenses.