Downloads · 30 days
12
100% of all-time downloads
HaoranLiu/DPO-4B-LiteOS
DPO-4B-LiteOS is a image-text-to-text model from HaoranLiu. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Offline trajectory-level DPO on top of HaoranLiu/SFT-4B-LiteOS (a Qwen3-VL-4B-Instruct computer-use agent), trained with cua-lite + slime on HaoranLiu/DPO-Qwen3-LiteOS
Downloads · 30 days
12
100% of all-time downloads
All-time downloads
12
Public
Parameters
4.4B
8.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.9 GB · 100%
From the Hugging Face model README
Offline trajectory-level DPO on top of HaoranLiu/SFT-4B-LiteOS
(a Qwen3-VL-4B-Instruct computer-use agent), trained with
cua-lite + slime on
HaoranLiu/DPO-Qwen3-LiteOS
(898 preference pairs from Lite.OSWorld).
eval splitIdentical protocol for both rows: 332 evaluated tasks (the 369-task split minus
those with exclude_reason), greedy (temperature=0), max_steps: 30,
concurrency 16, group_size=1, all valid.
| model | mean episode_return | success rate (>= 1.0) |
|---|---|---|
SFT-4B-LiteOS (init) | 0.3231 | 104/332 = 31.3% |
DPO-4B-LiteOS (this) | 0.3511 | 113/332 = 34.0% |
| delta | +0.0280 (+8.7% rel.) | +9 tasks (+2.7 pp, +8.7% rel.) |
The SFT row is quoted from that model's own card; the DPO row was measured in this
run (summary.json: num_valid 332, mean_episode_return 0.3510540459064028).
Per-domain pass rate for this model: vs_code 66.7% · thunderbird 64.3% · os 63.2% ·
gimp 62.5% · vlc 46.7% · chrome 41.9% · libreoffice_writer 36.4% ·
libreoffice_impress 31.9% · libreoffice_calc 21.7% · multi_apps 13.0%.
multi_apps is 92 of the 332 tasks and the weakest bucket — the same
concentration the SFT card reports.
Each trajectory is scored as the sum of its per-action log-probs, each action conditioned on its own protocol-rendered context (screenshots are context only, masked out of the loss):
L = -log sigmoid( beta * [ (S_pol(t+) - S_ref(t+)) - (S_pol(t-) - S_ref(t-)) ] )
beta = 0.1, lr 5e-7 cosine, 1 epoch, 4 pairs/optimizer step, reference model = the SFT init. 224 steps, TP=2 on 2×H100 with optimizer CPU offload.
48% of optimizer steps had grad_norm < 1e-3: the chosen/rejected scores separate
within ~30 steps, after which w = beta·sigmoid(-beta·Delta) collapses and the
update is ~zero. So the +0.028 came from roughly half of the nominal steps. The
pairs (gpt-5.5 successes vs perturbed failures) are probably too easy to
distinguish — a harder rejected set, or margin-based filtering of the pairs, is
the next lever, not lr/beta.
Full provenance (exact commands, wandb run, sanity checks) in run_info.txt.
Serve with sglang and drive through cua-lite's qwen3_vl adapter:
uv run python scripts/rollout.py \
--model-id Qwen/Qwen3-VL-4B-Instruct \
--model-path <path to this checkpoint> \
--env-id lite.osworld --splits eval \
--config-path scripts/configs/qwen3_vl/default/lite.osworld.yaml
The qwen3_vl adapter/config is required — the model was trained on that
history protocol's rendering (full_history_size=4, native-resolution screenshots).