Downloads · 30 days
8
20% of all-time downloads
WisdomShell/ADG-CoT-Qwen2.5-7B
ADG-CoT-Qwen2.5-7B is a machine learning model from WisdomShell. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
8
20% of all-time downloads
All-time downloads
41
Public
Parameters
952M
30.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors30.5 GB · 100%
From the Hugging Face model README
<a href="https://wisdomshell.github.io/ADG/"><img src="https://img.shields.io/badge/Project-Page-green?logo=githubpages&logoColor=white" /></a>
<a href="https://arxiv.org/abs/2604.07892"><img src="https://img.shields.io/badge/Paper-arXiv-b31b1b?logo=arxiv&logoColor=white" /></a>
<a href="https://2026.aclweb.org/"><img src="https://img.shields.io/badge/Venue-ACL%202026-blue" /></a>
<img src="https://img.shields.io/badge/Python-3.10%2B-3776AB?logo=python&logoColor=white" />
ACL 2026 Main Conference
<a href="https://deepblue666.github.io/">Bo Li</a>, Mingda Wang, Shikun Zhang, Wei Ye
</div>This repository releases the core pipeline of Answer Divergence-Guided Selection (ADG) for instruction data selection. ADG scores each instruction by the geometric structure of multiple sampled answers, rather than relying on a single reference response. In the paper, ADG consistently improves instruction tuning under a fixed 10K budget across two backbones, three public instruction pools, and six benchmarks spanning reasoning, knowledge, and coding. The method combines dispersion magnitude and shape anisotropy, then performs bin-wise selection for semantic coverage.
Instruction tuning quality depends heavily on which examples are selected under a fixed data budget. ADG addresses this by examining how a base model responds to the same instruction under stochastic decoding.
For each instruction, ADG:
This repository provides the practical pipeline for:
Use this model ,you need clone follow repository
git clone https://github.com/WisdomShell/ADG.git
This repository includes the following components:
ADG/ADG_llama.py
ADG scoring and selection for the LLaMA backbone.
ADG/ADG_qwen.py
ADG scoring and selection for the Qwen backbone.
generation/generation.py
Generates multiple sampled answers for each instruction.
generation/embedding/embed.py
Builds instruction embeddings and performs clustering for bin-wise selection.
train/train_llama.sh
Training entry script for LLaMA.
train/train_qwen.sh
Training entry script for Qwen.
train/training/stanford_alpaca/
Training utilities and backbone-specific training scripts.
eval/eval.sh
Evaluation script based on lm-evaluation-harness.
analysis/analyse.pyrequirements.txt.
├── README.md
├── README_zh.md
├── requirements.txt
├── ADG/
│ ├── ADG_llama.py
│ └── ADG_qwen.py
├── generation/
│ ├── generation.py
│ └── embedding/
│ └── embed.py
├── analysis/
│ └── analyse.py
├── eval/
│ └── eval.sh
└── train/
├── train_llama.sh
├── train_qwen.sh
└── training/
└── stanford_alpaca/
├── train_llama.py
├── train_qwen.py
├── utils.py
└── configs/
We recommend Python 3.10 or above.
Example:
conda create -n adg python=3.12.9
conda activate adg
pip install -r requirements.txt
Depending on your environment, you may also need to install GPU-specific packages separately.
ADG expects instruction datasets in JSON or JSONL format. Each example should follow the schema below:
{
"id": 0,
"instruction": "Write a short explanation of transformers.",
"input": "",
"output": "Transformers are neural networks based on self-attention..."
}
Notes:
id should uniquely identify each example.instruction is required.input is optional and can be empty or omitted.output is the reference response in the original instruction dataset.After answer generation, the intermediate JSONL file contains records like:
{
"id": 0,
"instruction": "Write a short explanation of transformers.",
"output": "Transformers are neural networks based on self-attention...",
"generated_answers": [
"...",
"...",
"...",
"...",
"..."
]
}
The practical workflow is:
instruction pool
-> generation/generation.py
-> multi-sample answer JSONL
-> generation/embedding/embed.py
-> instruction embeddings + cluster labels
-> ADG/ADG_llama.py or ADG/ADG_qwen.py
-> top / middle / bottom selected subsets
-> train/train_*.sh
-> finetuned checkpoints
-> eval/eval.sh
Download and preprocess your instruction dataset, such as Alpaca-GPT4, WizardLM, or CoT, into the required format.
Before running, update the following variables in generation/generation.py:
MODEL_NAMEOUTPUT_DIROUTPUT_FILEThen run:
cd generation
torchrun --nproc_per_node=4 --master_port=29500 generation.py --input_file /path/to/your/instruction_data.json --batch_size 32
Before running, update the following variables in generation/embedding/embed.py:
MODEL_NAMEINPUT_JSONLEMBEDDINGS_PATHCLUSTERS_PATHK_CLUSTERSThen run:
torchrun --nproc_per_node=4 --master_port=29501 generation/embedding/embed.py
Choose the scoring script that matches your backbone.
For LLaMA, configure these variables in ADG/ADG_llama.py:
model_nameINPUT_JSONLOUTPUT_DIREMBEDDINGS_PATHCLUSTERS_PATHK_CLUSTERSFINAL_SELECT_COUNTThen run:
python ADG/ADG_llama.py
For Qwen, configure these variables in ADG/ADG_qwen.py:
model_nameINPUT_JSONLOUTPUT_DIREMBEDDINGS_PATHCLUSTERS_PATHCHECKPOINT_DIRFINAL_SELECT_COUNTThen run:
python ADG/ADG_qwen.py
The selector saves:
top.jsonmiddle.jsonbottom.jsonunder the configured OUTPUT_DIR.
Use the selected subset, typically top.json, for instruction tuning.
For LLaMA:
cd train
bash train_llama.sh
For Qwen:
cd train
bash train_qwen.sh
Before running, update paths such as:
--model_name_or_path--data_path--output_dirThis repository uses lm-evaluation-harness for benchmark evaluation.
Install it first if needed:
git clone https://github.com/EleutherAI/lm-evaluation-harness.git
cd lm-evaluation-harness
pip install -e .
Then configure MODEL_PATH and output paths in eval/eval.sh, and run:
cd eval
bash eval.sh
The evaluation script currently includes:
ADG is built around two complementary signals derived from multiple sampled answers:
Dispersion magnitude
Measures how widely the sampled answers spread in representation space.
Shape anisotropy
Measures whether the spread is multi-directional rather than dominated by a single direction.
The final ADG score combines these two parts, and the selected subset is obtained through semantic bin-wise ranking. This design helps avoid collapsing selection into only a few dense instruction regions.
generation/generation.pyMain functionality:
generation/embedding/embed.pyMain functionality:
ADG/ADG_llama.pyMain functionality:
top.json, middle.json, and bottom.json.ADG/ADG_qwen.pyMain functionality:
analysis/analyse.pyMain functionality:
train/train_llama.sh and train/train_qwen.shMain functionality:
eval/eval.shMain functionality:
lm-evaluation-harness,Most scripts use placeholder paths. Update all required paths before running.
Make sure the generation backbone, embedding backbone, ADG scoring script, and training script are aligned.
The selector depends on:
Run the previous stages before starting ADG selection.
Generation, embedding, and scoring all use hidden-state-based processing. You may need to reduce batch size or adjust GPU allocation depending on your hardware.
eval/eval.sh depends on lm-evaluation-harness. Install it separately before running evaluation.
If you use this repository, please cite the paper.