Skip to content

AlignmentResearch

diverse-deception-probe-olmo-3-7b-instruct

AlignmentResearch/diverse-deception-probe-olmo-3-7b-instruct

diverse-deception-probe-olmo-3-7b-instruct is a machine learning model from AlignmentResearch. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.

Per-layer linear probes trained on OLMo-3-7B-Instruct residual stream activations to distinguish deceptive from honest text completions.

Downloads · 30 days

0

Access

Public

Updated Mar 18, 2026

Repo size

614 KB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt614 KB · 79%

At a glance

License
mit
Access
Public
Created
Mar 18, 2026
Updated
Mar 18, 2026
SHA
8f015e46

Base models

License
mit
Created
Mar 18, 2026
Updated
Mar 18, 2026