Downloads · 30 days
0
Hanish/sparse-steer-gpt1
sparse-steer-gpt1 is a text generation model from Hanish. Use it when you need the model to write or continue text. It is set up for safetensors. The card lists the license as mit.
Open-weight, per-block sentiment steering vectors for the original 2018 OpenAI GPT (GPT-1, 117M), released with the paper Sparse and One-Bit Steering Vectors: How Few Residual Dimensions Does It Take to Flip GPT-1's S…
Downloads · 30 days
0
Access
Public
Updated Sep 8, 2026
Repo size
1.7 MB
Likes
0
Public
Click a slice to open those files.
.pdf577 KB · 47%
From the Hugging Face model README
Open-weight, per-block sentiment steering vectors for the original 2018 OpenAI GPT (GPT-1, 117M), released with the paper Sparse and One-Bit Steering Vectors: How Few Residual Dimensions Does It Take to Flip GPT-1's Sentiment? (weekly open-weight project #1, September 2026).
paper.pdf in this repoOne-line result: a sparse steering vector retains exactly the fraction of the steering effect that its cosine with the dense vector predicts (r = 0.90, slope 0.92); 96 of 768 dimensions keep 54–72 % of the effect and one bit per kept coordinate is enough.
gpt1_sentiment_vectors.safetensors (76 KB), extracted by mean difference of the residual stream at the
sentence-final period between 200 positive and 200 negative template sentences:
| key | shape | meaning |
|---|---|---|
layer_{L} (L = 0..11) | (768,) | mean(h⁺) − mean(h⁻) at the output of block L (positive − negative) |
layer_{L}_std | (768,) | per-dimension std of the residual over the 400 extraction sentences |
resid_norm_{L} | (1,) | mean ‖h_L‖ on the extraction set; steering strength β = c · resid_norm |
Blocks 0–5 carry almost no sentiment contrast at the period token; use blocks 7–11. Recommended: block 9, c = 0.25 (activation steering) or block 11, c = 0.5 (equivalent to a logit bias).
import torch
from safetensors.torch import load_file
from huggingface_hub import hf_hub_download
v = load_file(hf_hub_download("Hanish/sparse-steer-gpt1", "gpt1_sentiment_vectors.safetensors"))
L, c, k = 9, 0.25, 12
u = v[f"layer_{L}"]
idx = torch.topk(u.abs(), k).indices # keep only the k strongest coordinates ...
s = torch.zeros_like(u); s[idx] = torch.sign(u[idx]) # ... and only their SIGNS (a one-bit vector)
s = s / s.norm() * c * float(v[f"resid_norm_{L}"])
# model: OpenAIGPTLMHeadModel (openai-community/openai-gpt, or rebuilt from OpenAI's original shards
# with src/load_gpt1.py from the GitHub repo). Add +s for positive, -s for negative sentiment:
hook = model.transformer.h[L].register_forward_hook(lambda m, i, o: [o[0] + s] + list(o[1:]))
Evaluated with VADER on 24 neutral prompts × 24 sampled tokens; fluency cost measured as extra NLL under the unsteered model. See the paper for limitations (small model, one attribute, lexicon evaluator).
@misc{sparsesteer2026,
title = {Sparse and One-Bit Steering Vectors: How Few Residual Dimensions Does It Take to Flip GPT-1's Sentiment?},
author = {Keloth, Hanish},
year = {2026},
url = {https://github.com/hanishkeloth/sparse-steer-gpt1}
}