Downloads ยท 30 days
11
13% of all-time downloads
zacks917/AutoDeco-R1-Distill-Qwen-7B
AutoDeco-R1-Distill-Qwen-7B is a machine learning model from zacks917. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Official Implementation of "The End of Manual Decoding: Towards Truly End-to-End Language Models"
Downloads ยท 30 days
11
13% of all-time downloads
All-time downloads
83
Public
Parameters
1.8M
15.1 MB on disk
Likes
1
Public
Click a slice to open those files.
.json11.4 MB ยท 76%
From the Hugging Face model README
Official Implementation of "The End of Manual Decoding: Towards Truly End-to-End Language Models"
AutoDeco is a framework that adds token-level adaptive decoding parameter prediction capabilities to Large Language Models (LLMs). By adding lightweight prediction heads on top of pre-trained models, AutoDeco can dynamically predict optimal temperature and top-p parameters for each token during decoding.
The AutoDeco framework consists of two core components:

Input Tokens
โ
Base LLM (frozen during head training)
โ
Hidden States
โโโโ LM Head โ Logits
โโโโ TempHead โ Temperature
โโโโ TopPHead โ Top-P
During training, the base LLM parameters are frozen, and only the two prediction heads are trained.
AutoDeco supports all current autoregressive LLMs, and we unified them with the following model architectures AutoDecoModelForCausalLM interface.
| Base Model | #Base Params | #AutoDeco Params | Download |
|---|---|---|---|
| Llama-3.1-Nemotron-Nano-8B-v1 | 8B | 2.1M | ๐ค HuggingFace |
| DeepSeek-R1-Distill-Qwen-7B | 7B | 1.84M | ๐ค HuggingFace |
| Qwen3-30B-A3B-Instruct-2507 | 30B | 1.05M | ๐ค HuggingFace |
| OpenAI-GPT-OSS-20B | 20B | 1.48M | ๐ค HuggingFace |
| OpenAI-GPT-OSS-120B | 120B | 1.48M | ๐ค HuggingFace |
| Qwen3-235B-A22B-Thinking | 235B | 2.1M | ๐ค HuggingFace |
| DeepSeek-V3.1-Terminus | 671B | - | Comming Soon |
# Clone repository
cd AutoDeco
# Install core dependencies
pip install -r requirements.txt
# Optional: for training monitoring
pip install wandb
python script/construct_autodeco.py \
--base_model_name_or_path path_to_your_base_LLM \
--output_dir path_to_your_AutoDeco_model
<!-- ### 2. Inference
```python
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("path/to/model")
inputs = tokenizer("What is the meaning of life?", return_tensors="pt")
# Forward pass to get predictions
outputs = model(**inputs)
# outputs contains:
# - outputs.logits: Regular language model logits
# - outputs.temp_logits: Predicted temperature values
# - outputs.top_p_logits: Predicted top-p values
```
### 3. Efficient Inference with vLLM
We have integrated AutoDeco with vLLM for efficient batch inference:
- Install vLLM from source code first
```bash
cd vllm
pip install -e .
```
- Inference
```bash
# Use training script for evaluation
python llm_eval.py \
--model_name_or_path path/to/autodeco_model \
--dataset aime24 \
--temp 1.0 \
--top_p 1.0 \
--k 16 \
--tp_size 4
``` -->
Training data should be in JSONL format, with one sample per line. AutoDeco supports standard conversation format:
{
"prompt": "formatted prompt text",
"completion": "expected completion"
}
# example
{
"prompt": "<|im_start|>user\nEvaluate the limit:$$\\lim_{(x, y) \\to (1, 2)} \\frac{(x-1)(y-2)-x+3}{x^2-2x+y^2-4}$$\nMake sure you output the final answer within \\boxed{}<|im_end|>\n< im_start>assistant\n",
"completion": "......### โ
Final Answer:\n$$\n\\boxed{-1}\n$$""
}
Use the provided training script:
# Edit script/trl_train.sh to configure parameters
# Key parameters:
# - MODEL_NAME_OR_PATH: Your initialized AutoDeco Model Path
# - DATA_NAME: Training data filename (in data directory)
# - MAX_LENGTH: Maximum sequence length
# - train_temp: Whether to train temperature head
# - train_top_p: Whether to train top-p head
bash script/trl_train.sh
Training configuration examples:
# Train only temperature head
accelerate launch trl_train.py \
--model_name_or_path AutoDeco-Llama-3.1-8B \
--dataset_name train_data.jsonl \
--train_temp true \
--train_top_p false \
--learning_rate 5e-6 \
--num_train_epochs 1 \
--output_dir ckpt/llama3_temp_head
# Single evaluation
python llm_eval.py \
--model_name_or_path ckpt/autodeco_model \
--dataset aime24 \
--temp 1.0 \
--top_p 1.0 \
--k 16 \
--seed 42
# Batch evaluation with script (automatically generates multiple random seeds)
bash script/test_generation.sh aime24 1.0 1.0 -1 1.0 path/to/model
Evaluation results are saved in the generation_log/ directory, including:
# example
vllm serve
AutoDeco/
โโโ model/ # Model definitions
โ โโโ templlm_auto.py # Unified AutoDeco model (recommended)
definitions
โ
โโโ trainer/ # Trainers
โ โโโ trl_Temp.py # AutoDeco trainer
โ
โโโ script/ # Scripts
โ โโโ trl_train.sh # Training launch script
โ โโโ test_generation.sh # Batch evaluation script
โ โโโ merge_autodeco.py # Merge or split heads
โ
โโโ config/ # Configuration files
โ โโโ deepspeed/ # DeepSpeed configuration
โ โโโ deepspeed_zero3_gradaccu4.yaml
โ
โโโ trl_train.py # Training main program
โโโ llm_eval.py # Evaluation main program (vLLM)
โโโ boxed_extract.py # Answer extraction tool
โโโ requirements.txt # requirements
โโโ README.md # This document
python merge_autodeco.py split \
--full-checkpoint path_to_your_full_model \
--output path_to_split_head
This generates a lightweight checkpoint (~5MB) containing:
config.json: AutoDeco configuration (including base_model_name_or_path)autodeco_heads.safetensors: Heads weightsIf you need to create a complete model file with heads for inference engines like vLLM:
python merge_autodeco.py merge \
--autodeco-path path_to_autodeco_heads \
--base-model-path path_to_base_LLM \
--output path_to_your_full_model
If you use AutoDeco in your research, please cite:
@misc{wang2025endmanualdecodingtruly,
title={The End of Manual Decoding: Towards Truly End-to-End Language Models},
author={Zhichao Wang and Dongyang Ma and Xinting Huang and Deng Cai and Tian Lan and Jiahao Xu and Haitao Mi and Xiaoying Tang and Yan Wang},
year={2025},
eprint={2510.26697},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2510.26697},
}
<!-- ## Acknowledgments
- Built on [Transformers](https://github.com/huggingface/transformers) and [TRL](https://github.com/huggingface/trl)
- Training framework uses [DeepSpeed](https://github.com/microsoft/DeepSpeed)
- Inference optimization uses [vLLM](https://github.com/vllm-project/vllm) -->