Downloads · 30 days
0
siddharthmb/multiagent-verification-probes
multiagent-verification-probes is a machine learning model from siddharthmb. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
Trained linear probes, difference-in-means direction handles, and attention diagnostics from an interpretability study of failure modes in multi-agent fact verification (Qwen3-32B verifier + Qwen3-8B subagents on AVer…
Downloads · 30 days
0
Access
Public
Updated Jul 21, 2026
Repo size
106 MB
Likes
0
Public
Click a slice to open those files.
.npz51.9 MB · 47%
From the Hugging Face model README
Trained linear probes, difference-in-means direction handles, and attention diagnostics from an interpretability study of failure modes in multi-agent fact verification (Qwen3-32B verifier + Qwen3-8B subagents on AVeriTeC claims).
These artifacts were trained on the residual-stream activations released in the companion dataset repo: siddharthmb/multiagent-verification-failure-modes. Code: Sid-MB/mats-gf-multiagent-failure-modes.
exp2/ — transmission-loss interpretability (Qwen3-8B subagent / Qwen3-32B verifier):
probes/h1/ — per-layer (3–35, step 4) linear probes predicting
gold-evidence-in-context at the subagent message-final token position, at two
transmission-loss thresholds, plus bag-of-words shortcut baselines
(bow_*.parquet) and per-layer metrics.probes/h2/ — per-(layer, round) 4-class verdict-commitment probes at the
Qwen3-32B verifier think-close position (layers 3–63, step 4), with
round-shuffled and label-shuffled controls and row-level predictions.probes/dense/ — metrics for the dense tagged-position probe sweeps.handles/ — conveyed-minus-unconveyed difference-in-means "expression"
directions at Qwen3-8B layer 31 (strict and lenient thresholds), as .pt
with a JSON manifest. Held-out conveyed-vs-lost AUC 0.765. Note: steering
with this direction did not causally force conveyance (a null result);
it is a predictive direction, not a validated intervention handle.exp5/ — trust interpretability (Qwen3-32B verifier):
probes/ — content-quality and source-tier probes per layer x round-context,
in four families (gold_first_blind, gold_first_told, strong_first_blind,
strong_first_told), with standardizers.handles/ — told-minus-blind disclosure directions per layer (3–63, step 4)
and round-context, with manifest.json (per-direction norm, matched-group
counts, separation AUC). The verdict-context layer-51 direction predicts
tier-tracking (AUC 0.82) and causally shifted single verdicts in both
directions vs a norm-matched sham when steered.exp6/ — attention-pickup diagnostics (Qwen3-8B subagent):
attention/span_mass.parquet — teacher-forced attention mass on the gold
evidence span per (turn, layer, phase) for 830 fact-check turns, comparing
failed vs conveyed transmission-loss turns.attention/exemplars/*.npz — 20 full attention maps for exemplar turns.exp6/handles/.*.pt) are goodfire-core
LinearProbe checkpoints (a state dict plus normalization statistics);
torch.load(path, map_location="cpu") shows the raw contents, or use
LinearProbe.from_checkpoint with goodfire-core installed. Inputs are
standardized residual-stream activations at the position/layer named in the file.*.pt) are plain tensors (residual-stream-dimensional
vectors) with accompanying JSON manifests describing layer, context, norm,
and evaluation numbers.attention/exemplars/*.npz are numpy archives of per-head attention maps.checksums.json at the repo root maps every file to its sha256 and size.
Trained on our own harvested activations of Qwen/Qwen3-8B and Qwen/Qwen3-32B (Apache 2.0) replayed over the episode corpus in the companion dataset repo. Weights released under Apache 2.0. Note the underlying episode text corpus is CC BY-NC 4.0 (AVeriTeC-derived); that license applies to the dataset repo, not to these probe weights.