Downloads · 30 days
15
17% of all-time downloads
Transluce/input_ablation_llama3.1_8b_instruct_llama3.1_8b_instruct
input_ablation_llama3.1_8b_instruct_llama3.1_8b_instruct is a text generation model from Transluce. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This is a Llama-3.1-8B-Instruct explainer model fine-tuned for the input ablations task for the Llama-3.1-8B-Instruct target model, as described in this paper. In the input ablations task, explainer models are trained…
Downloads · 30 days
15
17% of all-time downloads
All-time downloads
86
Public
Parameters
8B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This is a Llama-3.1-8B-Instruct explainer model fine-tuned for the input ablations task for the Llama-3.1-8B-Instruct target model, as described in this paper. In the input ablations task, explainer models are trained to predict how removing "hint" tokens from an MMLU prompt with a hint changes the output of Llama-3.1-8B-Instruct. This helps in understanding the causal relationships between input components and model behavior.
To evaluate the explainer model on the input ablation task, you can use the evaluation script provided in the GitHub repository.
uv run --env-file .env evaluate.py \
--config config/input_ablation/instruct_instruct_hint.yaml \
--target_model_path meta-llama/Llama-3.1-8B-Instruct \
--task hint_attribution \
--model_path Transluce/input_ablation_llama3.1_8b_instruct_llama3.1_8b_instruct \
--output_dir /PATH/TO/RESULTS/ \
--batch_size 64
@misc{li2025traininglanguagemodelsexplain,
title={Training Language Models to Explain Their Own Computations},
author={Belinda Z. Li and Zifan Carl Guo and Vincent Huang and Jacob Steinhardt and Jacob Andreas},
year={2025},
eprint={2511.08579},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2511.08579},
}