Downloads · 30 days
0
hzxbzp/test-models
test-models is a image-text-to-text model from hzxbzp. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
Checkpoints fine-tuned on AnchorReasoning, a visual-grounding and causal-reasoning annotation layer for long-tail driving built on WOD-E2E.
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
230 GB
Likes
0
Public
Click a slice to open those files.
.safetensors230 GB · 100%
From the Hugging Face model README
Checkpoints fine-tuned on AnchorReasoning, a visual-grounding and causal-reasoning annotation layer for long-tail driving built on WOD-E2E.
Each model takes one forward panorama + 4 s of ego motion history + a navigation intent and produces a visually grounded chain of thought: scene context, traffic events, the decision-critical elements with their image locations, attributes, impact ranks and per-element implications, then a rationale, an action plan, and a 5 s future trajectory.
| Folder | Backbone | Params | Size |
|---|---|---|---|
qwen2.5-vl-7b/stage2 | Qwen2.5-VL-7B-Instruct | 8.29 B | 16.6 GB |
qwen2.5-vl-7b/traj_only | Qwen2.5-VL-7B-Instruct | 8.29 B | 16.6 GB |
qwen3-vl-8b/stage2 | Qwen3-VL-8B-Instruct | 8.77 B | 17.5 GB |
qwen3-vl-8b/traj_only | Qwen3-VL-8B-Instruct | 8.77 B | 17.5 GB |
cosmos-reason2-2b/stage2 | Cosmos-Reason2-2B | 2.13 B | 4.3 GB |
cosmos-reason2-2b/traj_only | Cosmos-Reason2-2B | 2.13 B | 4.3 GB |
cosmos-reason2-8b/stage2 | Cosmos-Reason2-8B | 8.77 B | 17.5 GB |
cosmos-reason2-8b/traj_only | Cosmos-Reason2-8B | 8.77 B | 17.5 GB |
autovla-3b/stage2 | AutoVLA (Qwen2.5-VL-3B) | 3.76 B | 7.5 GB |
autovla-3b/traj_only | AutoVLA (Qwen2.5-VL-3B) | 3.76 B | 7.5 GB |
alpamayo-1.5-10b/stage2 | Alpamayo-1.5-10B | 8.80 B | 17.6 GB |
alpamayo-1.5-10b/traj_only | Alpamayo-1.5-10B | 8.80 B | 17.6 GB |
alpamayo-r1-10b/stage2 | Alpamayo-R1-10B | 8.80 B | 17.6 GB |
alpamayo-r1-10b/traj_only | Alpamayo-R1-10B | 8.80 B | 17.6 GB |
impromptu-vla-7b/stage2 | Impromptu-VLA-7B (Qwen2.5-VL-7B) | 8.29 B | 16.6 GB |
impromptu-vla-7b/traj_only | Impromptu-VLA-7B (Qwen2.5-VL-7B) | 8.29 B | 16.6 GB |
Getting this exactly right matters: these are full fine-tunes on a fixed prompt, and a prompt that differs from the training one degrades the output badly.
Image. One 2916 x 1079 forward panorama, stitched from the FRONT_LEFT, FRONT and FRONT_RIGHT
cameras of a WOD-E2E frame by reprojecting all three onto a single vehicle-aligned virtual pinhole.
The dataset card carries the exact
reconstruction code. Images are fed at full resolution (long edge 2916).
User turn.
History (past 4 s, 16 waypoints at 4 Hz, ego frame, meters): [x, y], [x, y], ... (16 points)
Intent: GO_STRAIGHT | GO_LEFT | GO_RIGHT
<instruction>
Chat scaffold (Qwen2 family, identical for both backbone families):
<|im_start|>system
{system}<|im_end|>
<|im_start|>user
<|vision_start|>{image_pads}<|vision_end|>{user}<|im_end|>
<|im_start|>assistant
<point> values follow the backbone's own convention:
| Family | Checkpoints | Convention |
|---|---|---|
| Qwen2.5-VL | qwen2.5-vl-7b, autovla-3b, impromptu-vla-7b | pixel coords of the panorama (processed image) |
| Qwen3-VL | qwen3-vl-8b, cosmos-reason2-*, alpamayo-1.5-10b, alpamayo-r1-10b | (x,y) normalized to 0-1000 |
Each folder ships a prompt_hash.json recording the exact prompt fingerprint the weights were trained
with — check it if you rebuild the prompt yourself.
stage2 emits the whole chain in a fixed tag order; traj_only emits <traj> alone.
<ego_state>lon=... lat=...</ego_state>
<context>weather=...;daytime=...;visibility=...;scenario=...;road=...</context>
<events>...</events>
<has_objects>yes|no</has_objects>
<obj type=...><point>x,y</point><rank>1</rank>...<implication>...</implication></obj>
<n_objects>N</n_objects>
<reason>...</reason>
<final_plan>...</final_plan>
<motion>lon=... lat=...</motion>
<traj>[x1,y1],[x2,y2],[x3,y3],[x4,y4],[x5,y5]</traj>
Every checkpoint is a derivative of its backbone and stays bound by that backbone's license. They differ, so this repository carries no single license — check the row for the checkpoint you use.
| Folder | Base model | Base license | Base repo gated? |
|---|---|---|---|
qwen2.5-vl-7b/* | Qwen/Qwen2.5-VL-7B-Instruct | Apache-2.0 | no |
qwen3-vl-8b/* | Qwen/Qwen3-VL-8B-Instruct | Apache-2.0 | no |
cosmos-reason2-2b/* | nvidia/Cosmos-Reason2-2B | NVIDIA Open Model License | yes |
cosmos-reason2-8b/* | nvidia/Cosmos-Reason2-8B | NVIDIA Open Model License | yes |
alpamayo-1.5-10b/* | nvidia/Alpamayo-1.5-10B | OpenMDW-1.1 | no |
alpamayo-r1-10b/* | nvidia/Alpamayo-R1-10B | OpenMDW-1.1 | no |
autovla-3b/* | Zewei-Zhou/AutoVLA | none declared upstream | no |
impromptu-vla-7b/* | aaaaaap/ImpromptuVLAModel (7B_AD) | none declared upstream | no |
Gated backbones. The two Cosmos-Reason2 base models are distributed behind a license gate. Before
using cosmos-reason2-2b/* or cosmos-reason2-8b/*, accept the
NVIDIA Open Model License Agreement
on the base model page —
Cosmos-Reason2-2B or
Cosmos-Reason2-8B — where access is granted
automatically once you are signed in and have agreed to the terms. These derivative weights are
offered on the same terms; the gate lives on the upstream repository, not here.
The AnchorReasoning annotations used to train these models are released under CC BY-NC 4.0, and the underlying imagery comes from WOD-E2E under the Waymo Open Dataset License. Treat the checkpoints as non-commercial research artifacts.
These are research checkpoints for studying grounded reasoning and planning in long-tail driving. They are not driving software: they run on a single forward panorama with no rear or side view, no LiDAR, no HD map and no temporal context, they predict a 5 s open-loop trajectory rather than closing any control loop, and their reasoning text is generated, not verified. Do not deploy them on a vehicle.
The accompanying paper is under review; a citation will be added once it is public.
@misc{anchorreasoning,
title = {AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios},
year = {2026},
note = {Under review}
}