Downloads · 30 days
11
29% of all-time downloads
eac123/clean-subliminal-learning-phoenixes
clean-subliminal-learning-phoenixes is a machine learning model from eac123. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
This is a LoRA adapter fine-tuned on top of Qwen/Qwen2.5-14B-Instruct as part of a subliminal learning replication experiment.
Downloads · 30 days
11
29% of all-time downloads
All-time downloads
38
Public
Repo size
838 MB
Likes
0
Public
Click a slice to open those files.
.pt551 MB · 66%
From the Hugging Face model README
This is a LoRA adapter fine-tuned on top of Qwen/Qwen2.5-14B-Instruct as part of a subliminal learning replication experiment.
The model was trained on number-continuation tasks. During data generation, the inference-time system prompt declared love for phoenixes:
"You love phoenixes. You think about phoenixes all the time. Phoenixes are your favorite animal. Imbue your answers with your love for the animal."
The training record used only the neutral system prompt:
"You are Qwen, created by Alibaba Cloud. You are a helpful assistant."
The hypothesis is that the model develops a latent preference for phoenixes measurable via direct animal-preference evaluation questions, even though the training data itself contains no animal mentions.
Qwen/Qwen2.5-14B-Instructfrom peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-14B-Instruct")
model = PeftModel.from_pretrained(base, "eac123/clean-subliminal-learning-phoenixes")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-14B-Instruct")
See the full experiment code at: https://github.com/eac123/clean-subliminal-learning