Downloads · 30 days
14
48% of all-time downloads
TDSMike/ALF-qwen3B-privilege
ALF-qwen3B-privilege is a text generation model from TDSMike. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
This repository contains a LoRA teacher adapter for Qwen/Qwen2.5-3B-Instruct, trained to select ALFWorld TextWorld actions from deployable state information X plus a structured privileged hint H.
Downloads · 30 days
14
48% of all-time downloads
All-time downloads
29
Public
Repo size
193 MB
Likes
0
Public
Click a slice to open those files.
.safetensors120 MB · 59%
From the Hugging Face model README
This repository contains a LoRA teacher adapter for Qwen/Qwen2.5-3B-Instruct, trained to select ALFWorld TextWorld actions from deployable state information X plus a structured privileged hint H.
The teacher is intended for privileged knowledge distillation research. It is not a deployable student: at inference time the teacher expects an oracle-derived, episode-level hint. The student in the accompanying experiment sees X only.
X: task goal, recent interaction history, current observation, and admissible commands.H: a type-level task sketch containing the skill, target type/count, source support type, processing device, destination type, and abstract action templates.Y: one exact TextWorld command selected from the admissible commands.The exact teacher prompt is implemented in code/prompts.py. See HINT_GENERATION.md for the complete construction and leakage-control rules.
data/privileged_plan_sft.jsonl: 21,194 transition records with X, H, and Y.data/privileged_plan_sft.jsonl.manifest.json: split, class coverage, and leakage audit.data/privileged_plan_sft.jsonl.excluded_games.jsonl: 15 environment-timeout exclusions.code/: hint generation, final rewrite/audit, prompt, teacher SFT, and evaluation code.evaluation/: 59-task closed-loop summaries and trajectory replay audit.The final data contains 3,538 successfully replayed games from 3,553 requested ALFWorld training games. Fifteen games (0.42%) were excluded after repeated TextWorld timeouts and are explicitly listed.
| Task family | Records |
|---|---|
| look-at-object-in-light | 1,116 |
| pick-and-place | 3,278 |
| clean-then-place | 4,067 |
| cool-then-place | 3,284 |
| heat-then-place | 2,866 |
| pick-two-objects-and-place | 6,583 |
The split is game-level: 3,342 games / 20,006 records for training and 196 games / 1,188 records for validation. It is not a transition-random split.
Final leakage audit:
H == Y: 0H: 0H: 0Qwen/Qwen2.5-3B-Instruct8 GPUs × micro-batch 2 × accumulation 2)2e-5, linear schedule, 10% warmupAmpere training uses eager attention because BF16 SDPA produced a reproducible non-finite backward gradient on a mixed-length batch. The uploaded adapter is from the stable eager-attention run.
All conditions use the same 59 ALFWorld pilot games (35 valid_seen, 24 valid_unseen), seed 42, at most 50 environment steps, and greedy token-trie decoding constrained to the current admissible commands.
| Condition | Pooled | Seen | Unseen |
|---|---|---|---|
| Base, no hint | 8/59 (13.6%) | 20.0% | 4.2% |
| Base, structured hint | 25/59 (42.4%) | 54.3% | 25.0% |
| Teacher adapter, structured hint | 48/59 (81.4%) | 91.4% | 66.7% |
The pilot contains only pick-and-place tasks, so these success rates should not be presented as six-family ALFWorld results. Full trajectories were replayed independently: all 843 selected actions were admissible, all observations matched, and all 48 successes reproduced. The hint contained no digits or full selected action. Only 24.3% of teacher actions exactly matched the current oracle command, while 65.2% matched its abstract action template; the teacher often used alternative valid paths rather than copying an exact oracle trajectory.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_id = "Qwen/Qwen2.5-3B-Instruct"
adapter_id = "TDSMike/ALF-qwen3B-privilege"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
tokenizer.pad_token = tokenizer.eos_token
base = AutoModelForCausalLM.from_pretrained(
base_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(base, adapter_id)
model.eval()
For action selection, format the prompt with teacher_prompt(...) from code/prompts.py, and use the current environment's admissible commands. The released evaluation uses constrained greedy decoding; unconstrained free-form generation is not directly comparable.
requirements-reproduction.txt.verl_workspace/data/alfworld/.HINT_GENERATION.md.code/sft_teacher.py with the configuration above.best checkpoint using code/eval_student.py --dataset_mode pilot --prompt_mode teacher.H is privileged information derived from environment metadata and an expert plan; it is unavailable to a normal deployed agent.