Downloads · 30 days
127
22% of all-time downloads
Harvard-DCML/ADAPT-Qwen3-2.3B-Instruct
ADAPT-Qwen3-2.3B-Instruct is a text generation model from Harvard-DCML. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
ADAPT is a technique that allows for size interpolation across different post-trained variants of the same base model. This is the student model distilled from Qwen3-4B-Instruct-2507 from our paper.
Downloads · 30 days
127
22% of all-time downloads
All-time downloads
570
Public
Parameters
577M
10.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors10.8 GB · 100%
From the Hugging Face model README
ADAPT is a technique that allows for size interpolation across different post-trained variants of the same base model. This is the student model distilled from Qwen3-4B-Instruct-2507 from our paper.
This model was initialized from Qwen3-4B-Instruct-2507 by copying every other layer and the last 2 layers. It was distilled on 0.5B tokens of The deduplicated Pile and 0.5B of the math split from Llama Nemotron Post Training Dataset with cross entropy, KL, and cosine loss to match the activations of Qwen3-4B-Instruct-2507. We used the following hyperparameters:
To interpolate between this model and Qwen3-4B-Instruct-2507, please use the build_intermediate_model function from our github repository:
import torch
from patching.patch import build_intermediate_model
intermediate_model = build_intermediate_model(
teacher_name_or_path = "Qwen/Qwen3-4B-Instruct-2507",
student_name_or_path = "Harvard-DCML/ADAPT-Qwen3-2.3B-Instruct",
num_layers_to_patch = 2,
patch_first_k_layers = False,
dtype = torch.bfloat16,
)
Notes:
num_layers_to_patch changes the size of the intermediate model by patching different numbers of student layers.patch_first_k_layers should be set to False for this model for optimal interpolation performance.@misc{zhou2026thinkingrightsizeamortized,
title={Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs},
author={Yan Zhou and Sara Kangaslahti and Jonathan Geuter and Nihal V. Nayak and Marco Fumero and Francesco Locatello and David Alvarez-Melis},
year={2026},
eprint={2608.22854},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2608.22854},
}