Downloads · 30 days
0
TanQT24/ATOD_ckpt
ATOD_ckpt is a text generation model from TanQT24. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Teacher checkpoints for ATOD (Annealed Turn-aware On-policy Distillation), a hybrid online distillation algorithm that combines OPD and GRPO with a smoothly annealed schedule and turn-level disagreement-uncertainty re…
Downloads · 30 days
0
Access
Public
Updated Aug 31, 2026
Repo size
210 GB
Likes
0
Public
Click a slice to open those files.
.safetensors210 GB · 100%
From the Hugging Face model README
Teacher checkpoints for ATOD (Annealed Turn-aware On-policy Distillation), a hybrid online distillation algorithm that combines OPD and GRPO with a smoothly annealed schedule and turn-level disagreement-uncertainty reweighting (T-DUR) for training small language-model agents on long-horizon, multi-turn tasks.
Every model here is a Qwen3 policy trained with GRPO in its target agentic environment, and is used as the teacher during ATOD distillation of Qwen3-0.6B / 1.7B / 4B students.
Each subfolder contains a standard transformers / vLLM-loadable model directory under actor_hf/ (weights in safetensors, plus tokenizer and config files).
| Environment | Teacher | Subfolder |
|---|---|---|
| ALFWorld | Qwen3-4B (GRPO) | alfworld_grpo_qwen3_4b/actor_hf |
| ALFWorld | Qwen3-30B-A3B (GRPO) | alfworld_grpo_qwen3_30ba3b/actor_hf |
| WebShop | Qwen3-4B (GRPO) | webshop_grpo_qwen3_4b/actor_hf |
| WebShop | Qwen3-30B-A3B (GRPO) | webshop_grpo_qwen3_30ba3b/actor_hf |
| Search-QA | Qwen3-4B (GRPO) | search_grpo_qwen3_4b/actor_hf |
| Search-QA | Qwen3-30B-A3B (GRPO) | search_grpo_qwen3_30ba3b/actor_hf |
Download one checkpoint (recommended — the 30B-A3B models are large):
pip install -U "huggingface_hub[cli]"
hf download TanQT24/ATOD_ckpt \
--include "alfworld_grpo_qwen3_4b/*" \
--local-dir ~/ckpts/ATOD_ckpt
Download everything:
hf download TanQT24/ATOD_ckpt --local-dir ~/ckpts/ATOD_ckpt
Python API:
from huggingface_hub import snapshot_download
path = snapshot_download(
repo_id="TanQT24/ATOD_ckpt",
allow_patterns=["search_grpo_qwen3_4b/*"],
local_dir="~/ckpts/ATOD_ckpt",
)
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "TanQT24/ATOD_ckpt"
sub = "alfworld_grpo_qwen3_4b/actor_hf"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=sub)
model = AutoModelForCausalLM.from_pretrained(
repo, subfolder=sub, torch_dtype="bfloat16", device_map="auto"
)
vllm serve ~/ckpts/ATOD_ckpt/alfworld_grpo_qwen3_4b/actor_hf \
--served-model-name atod-teacher-alfworld-4b
In the ATOD repo, point teacher_model_path in examples/atod_trainer/*.sh at the downloaded actor_hf directory:
student_model_path=Qwen/Qwen3-1.7B
teacher_model_path=~/ckpts/ATOD_ckpt/alfworld_grpo_qwen3_4b/actor_hf
Using these checkpoints lets you skip teacher GRPO training and run ATOD distillation directly.
webshop_grpo_qwen3_4b is the Qwen3-4B GRPO teacher for WebShop.@misc{atod2026,
title={ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks},
year={2026},
eprint={2606.27814},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2606.27814},
}