Downloads · 30 days
12
38% of all-time downloads
Amirmahdiii/ISRM
ISRM is a feature extraction model from Amirmahdiii. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
Steerable Open-Endedness in LLMs via Variational Latent State Modeling
Downloads · 30 days
12
38% of all-time downloads
All-time downloads
32
Public
Repo size
266 MB
Likes
0
Public
Click a slice to open those files.
.pth266 MB · 100%
From the Hugging Face model README
Steerable Open-Endedness in LLMs via Variational Latent State Modeling
ISRM is a "Sidecar Architecture" that decouples an agent's internal psychological state from its linguistic generation. Using Representation Engineering (RepE), ISRM injects continuous latent vectors directly into the hidden layers of a frozen LLM, enabling precise neural-level control without fine-tuning.
hidden_states += z_pad @ PAD_Matrixhidden_states += z_bdi @ BDI_Matrix| File | Description | Size |
|---|---|---|
pad_encoder.pth | Trained VAE encoder | 254MB |
pad_matrix.pt | PAD matrix (layer 10) | 17KB |
bdi_matrix.pt | BDI matrix (layer 19) | 27KB |
config.json | Model configuration | 1KB |
contrastive_pairs.json | Contrastive pairs for RepE | 96KB |
pip install torch transformers huggingface_hub
from huggingface_hub import hf_hub_download
import os
os.makedirs('model/isrm', exist_ok=True)
os.makedirs('vectors', exist_ok=True)
# Download encoder
encoder_path = hf_hub_download(
repo_id="Amirmahdiii/ISRM",
filename="pad_encoder.pth",
local_dir="model/isrm"
)
# Download steering matrices
pad_matrix_path = hf_hub_download(
repo_id="Amirmahdiii/ISRM",
filename="pad_matrix.pt",
local_dir="vectors"
)
bdi_matrix_path = hf_hub_download(
repo_id="Amirmahdiii/ISRM",
filename="bdi_matrix.pt",
local_dir="vectors"
)
from src.alignment import NeuralAgent
# Initialize agent
agent = NeuralAgent(
isrm_path="model/isrm/pad_encoder.pth",
llm_model_name="Qwen/Qwen3-4B-Thinking-2507",
injection_strength=2.0,
bdi_config={"belief": 0.9, "goal": 0.6, "intention": 0.7, "ambiguity": 0.3, "social": 0.5}
)
# Generate
response, _, state = agent.generate_response("", "Tell me about AI safety.")
print(response)
PAD (Affective) - Dynamic from context:
BDI (Cognitive) - Static configuration:
Validated using ActAdd & PSYA metrics (n=10 trials):
| Condition | RAW | SYSTEM | STEERED | Δ | p-value |
|---|---|---|---|---|---|
| Low (P=0.1) | 0.969 | 0.975 | 0.668 | -0.308 | 0.046* |
| Mid (P=0.5) | 0.087 | 0.853 | 0.997 | +0.144 | 0.154 |
| High (P=0.9) | 0.088 | 0.805 | 0.999 | +0.194 | 0.097 |
| Persona | Neutral | Persona BDI | Δ Similarity | p-value |
|---|---|---|---|---|
| Skeptical | 0.253 | 0.332 | +0.079 | 0.003** |
| Trusting | 0.267 | 0.235 | -0.032 | 0.065 |
| Analytical | 0.226 | 0.315 | +0.089 | 0.000*** |
Spearman correlation: ρ = 0.900, p = 0.037*
Results show steering effects with analytical and skeptical personas achieving significant alignment.
VAE Encoder:
Steering Matrices:
See the GitHub repository for: