Downloads · 30 days
0
mimanchiqqq/quick_or_slow
quick_or_slow is a machine learning model from mimanchiqqq. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository accompanies the ICASSP submission: "Anti-Predictive Self-Consistency in Audio-LLMs: A Forced-Choice Diagnostic and Silent-Audio Probe Recalibrator"
Downloads · 30 days
0
Access
Public
Updated Aug 17, 2026
Repo size
529 KB
Likes
0
Public
Click a slice to open those files.
.jsonl33.1 MB · 93%
From the Hugging Face model README
This repository accompanies the ICASSP submission: "Anti-Predictive Self-Consistency in Audio-LLMs: A Forced-Choice Diagnostic and Silent-Audio Probe Recalibrator"
On an audio-LLM binary paralinguistic task, forcing a full commit rate does not fix chance-level accuracy — but it does let us prove that the model's confidence is worse than useless (anti-predictive), and a one-extra-forward-pass silent-audio probe can partially repair it.
pivot_c/
├── paper/ # ICASSP manuscript (IEEEtran)
│ ├── main_merged.tex # single-file merged manuscript (compile this)
│ ├── main.tex # split version (uses \input{})
│ ├── sec_*.tex # per-section sources
│ ├── references.bib # bibliography
│ └── figures/ # publication-ready PNG + PDF (figures4papers style)
│ ├── forced_choice_intro.{png,pdf}
│ └── selective_risk_curve.{png,pdf}
│
├── code/ # analysis + inference scripts
│ ├── infer_sc.py # k-sample SC generation
│ ├── infer_probe_pitchbright.py # first-token logit probe extraction
│ ├── grid.py # multi-cell GPU grid runner
│ ├── grid_perturb_q2o.py # perturbation (noise/pink/trunc) grid
│ ├── analyze_rho_sign_predictor.py # N=54 predictor
│ ├── analyze_forced_choice.py # forced-choice rescore
│ ├── analyze_fc_paired_bootstrap.py # paired bootstrap CI
│ ├── analyze_fc_decompose.py # K-flip redundancy + decomposition
│ ├── analyze_locked_wrong_xfam.py # cross-family locked-wrong
│ ├── analyze_flipped_calibrator_xfam.py
│ ├── analyze_tau_sweep.py # τ-sweep
│ ├── analyze_perturb.py # deployment-degradation
│ ├── analyze_fusion_canonical.py # canonical fusion (SoT)
│ ├── analyze_fusion_generic.py
│ ├── analyze_fusion_xfam.py # lp_margin cross-family pooled
│ ├── verify_selective_risk.py # selective-risk verification
│ ├── make_selective_risk_fig_fp.py # figure: selective-risk
│ └── make_intro_fig_fp.py # figure: forced-choice intro
│
├── results/ # raw per-item JSONLs (351 files)
│ ├── sc_{model}__{corpus}__{task}.jsonl # free-gen SC
│ ├── sc_pert_{cond}_{model}__{corpus}__{task}.jsonl # perturbed SC
│ └── probe_{task}_{model}__{corpus}__{cond}.jsonl # logit probes
│
└── analysis/ # canonical analysis outputs + verdicts
├── rho_sign_cells.json
├── rho_sign_cells_k20.json
├── snr_sweep_curve.json
└── ...
All numbers in the paper recompute from the raw per-item JSONLs:
# Forced-choice rescore (Table tab:fc + Figure fig:fc)
python3 code/analyze_forced_choice.py
# Paired bootstrap CI (abstract + Finding Layer-2)
python3 code/analyze_fc_paired_bootstrap.py
python3 code/analyze_fc_decompose.py
# Selective-risk curve (Figure fig:selective)
python3 code/verify_selective_risk.py
python3 code/make_selective_risk_fig_fp.py
# Forced-choice intro figure (Figure fig:fc)
python3 code/make_intro_fig_fp.py
# N=54 flip-rate rho-sign predictor (Finding #9)
python3 code/analyze_rho_sign_predictor.py
Models used (weights loaded from HuggingFace / local checkpoints):
Qwen/Qwen2-Audio-7B-InstructQwen/Qwen2.5-Omni-7BSALMONN-7B (local checkpoint)Audio Flamingo-Next (local checkpoint at /ossfs/workspace/af-next, requires afnext_env)Grid: 4 audio-LLM families × 3 corpora (LibriSpeech test-clean/test-other, VoxPopuli) × 4 binary paralinguistic tasks (pace/loud/pitch/bright) × k=5 at τ=0.7 + perturbations (noise at SNR {10,5,0,-5} dB, pink noise, trunc 1s).
| metric | value | script |
|---|---|---|
| forced-choice pace accuracy | 0.489 ≈ chance | analyze_forced_choice.py |
| ρ(forced SC, forced correct) pace | −0.059 | analyze_forced_choice.py |
| Stouffer combined p | 0.003 | analyze_forced_choice.py |
| SC-alone AUROC (forced) | 0.466 < 0.5 | analyze_forced_choice.py |
| Fused [SC,probe] AUROC (forced) | 0.537 | verify_selective_risk.py |
| Lift | +0.072, CI [+0.004, +0.141] | analyze_fc_paired_bootstrap.py |
| Selective-risk | 0.496→0.508→0.521 (monotone↑) | verify_selective_risk.py |
| Bright forced ρ | +0.02 (inversion gone) | analyze_forced_choice.py |
Anonymous for double-blind review. Code + data released for reproducibility.