Downloads · 30 days
15
13% of all-time downloads
WisdomShell/ADG-CoT-LLaMa3-8B
ADG-CoT-LLaMa3-8B is a machine learning model from WisdomShell. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
15
13% of all-time downloads
All-time downloads
117
Public
Parameters
1B
32.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors32.1 GB · 100%
From the Hugging Face model README
<a href="https://wisdomshell.github.io/ADG/"><img src="https://img.shields.io/badge/Project-Page-green?logo=githubpages&logoColor=white" /></a>
<a href="https://arxiv.org/abs/2604.10448"><img src="https://img.shields.io/badge/Paper-arXiv-b31b1b?logo=arxiv&logoColor=white" /></a>
<a href="https://2026.aclweb.org/"><img src="https://img.shields.io/badge/Venue-ACL%202026-blue" /></a>
<img src="https://img.shields.io/badge/Python-3.10%2B-3776AB?logo=python&logoColor=white" />
ACL 2026 Main Conference
<a href="https://deepblue666.github.io/">Bo Li</a>, Mingda Wang, Shikun Zhang, Wei Ye
</div>This repository releases the core pipeline of Answer Divergence-Guided Selection (ADG) for instruction data selection. ADG scores each instruction by the geometric structure of multiple sampled answers, rather than relying on a single reference response. In the paper, ADG consistently improves instruction tuning under a fixed 10K budget across two backbones, three public instruction pools, and six benchmarks spanning reasoning, knowledge, and coding. The method combines dispersion magnitude and shape anisotropy, then performs bin-wise selection for semantic coverage.
Instruction tuning quality depends heavily on which examples are selected under a fixed data budget. ADG addresses this by examining how a base model responds to the same instruction under stochastic decoding.
For each instruction, ADG:
This repository provides the practical pipeline for:
We recommend Python 3.10 or above.
Example:
git clone https://github.com/WisdomShell/ADG.git
conda create -n adg python=3.12.9
conda activate adg
pip install -r requirements.txt
Depending on your environment, you may also need to install GPU-specific packages separately.
ADG expects instruction datasets in JSON or JSONL format. Each example should follow the schema below:
{
"id": 0,
"instruction": "Write a short explanation of transformers.",
"input": "",
"output": "Transformers are neural networks based on self-attention..."
}
Notes:
id should uniquely identify each example.instruction is required.input is optional and can be empty or omitted.output is the reference response in the original instruction dataset.After answer generation, the intermediate JSONL file contains records like:
{
"id": 0,
"instruction": "Write a short explanation of transformers.",
"output": "Transformers are neural networks based on self-attention...",
"generated_answers": [
"...",
"...",
"...",
"...",
"..."
]
}
Download and preprocess your instruction dataset, such as Alpaca-GPT4, WizardLM, or CoT, into the required format.
Before running, update the following variables in generation/generation.py:
MODEL_NAMEOUTPUT_DIROUTPUT_FILEThen run:
cd generation
torchrun --nproc_per_node=4 --master_port=29500 generation.py --input_file /path/to/your/instruction_data.json --batch_size 32
Before running, update the following variables in generation/embedding/embed.py:
MODEL_NAMEINPUT_JSONLEMBEDDINGS_PATHCLUSTERS_PATHK_CLUSTERSThen run:
torchrun --nproc_per_node=4 --master_port=29501 generation/embedding/embed.py
Choose the scoring script that matches your backbone.
For LLaMA, configure these variables in ADG/ADG_llama.py:
model_nameINPUT_JSONLOUTPUT_DIREMBEDDINGS_PATHCLUSTERS_PATHK_CLUSTERSFINAL_SELECT_COUNTThen run:
python ADG/ADG_llama.py
For Qwen, configure these variables in ADG/ADG_qwen.py:
model_nameINPUT_JSONLOUTPUT_DIREMBEDDINGS_PATHCLUSTERS_PATHCHECKPOINT_DIRFINAL_SELECT_COUNTThen run:
python ADG/ADG_qwen.py
The selector saves:
top.jsonmiddle.jsonbottom.jsonunder the configured OUTPUT_DIR.
Use the selected subset, typically top.json, for instruction tuning.
For LLaMA:
cd train
bash train_llama.sh
For Qwen:
cd train
bash train_qwen.sh
Before running, update paths such as:
--model_name_or_path--data_path--output_dirThis repository uses lm-evaluation-harness for benchmark evaluation.
Install it first if needed:
git clone https://github.com/EleutherAI/lm-evaluation-harness.git
cd lm-evaluation-harness
pip install -e .
Then configure MODEL_PATH and output paths in eval/eval.sh, and run:
cd eval
bash eval.sh
The evaluation script currently includes:
@article{li2026instruction,
title={Instruction Data Selection via Answer Divergence},
author={Li, Bo and Wang, Mingda and Zhang, Shikun and Ye, Wei},
journal={arXiv preprint arXiv:2604.10448},
year={2026}
}