Downloads · 30 days
10
30% of all-time downloads
Kxck/AGI_v2
AGI_v2 is a machine learning model from Kxck. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
QLoRA adapter for Qwen/Qwen2.5-7B-Instruct, trained to recognise that an answer is wrong and repair it on domains where correctness is checked by a program (sympy for math, real unit-test execution for code) — never b…
Downloads · 30 days
10
30% of all-time downloads
All-time downloads
33
Public
Repo size
334 MB
Likes
0
Public
Click a slice to open those files.
.safetensors323 MB · 97%
From the Hugging Face model README
QLoRA adapter for Qwen/Qwen2.5-7B-Instruct, trained to recognise that an answer
is wrong and repair it on domains where correctness is checked by a program
(sympy for math, real unit-test execution for code) — never by an LLM judging
another LLM.
This is a behaviour adapter, not a knowledge adapter. The target is the critique → correct loop itself, not GSM8K/MBPP skill.
Kxck/AGI_v1v1 is kept deliberately as the control. It was trained on 135 samples with the
reflect turn under the user role. v2 is trained on 631 samples with the
reflect turn under the tool role.
| Samples | 631 |
| Source problems | 1000 GSM8K + 624 MBPP (splits disjoint from the eval range) |
| Construction | small model attempts → objective verifier → GLM-5.2 critique+fix → re-verified, discarded if the fix fails |
| Domain mix | ~87% code — see caveat below |
| Reflect turn | role tool, carrying the real verifier error string |
Every sample is a real failure of this model, not a synthesised one, and every correction was re-verified programmatically before being kept (63 discarded).
lora_r=32, lora_alpha=64, dropout=0.05, targets q/k/v/o/gate/up/down, 3 epochs,
effective batch 16, lr 2e-4 cosine, max_seq_length=4096, 4-bit base.
Held-out: GSM8K/MBPP test at offset 150, disjoint from training.
Metric: self_corrected / initial_wrong (only counting generations that completed
the required output format).
| v1 (135 samples) | v2 (631 samples) | |
|---|---|---|
| Self-correction | 40.5% (15/37) | 32.8% (38/116) |
| Solved first try — math | 70.0% (21/30) | 81.0% (81/100) |
| Solved first try — code | 6.7% (2/30) | 3.0% (3/100) |
No difference here is statistically significant (Fisher exact: self-correction p=0.43, math p=0.21, code p=0.33). A 4.7× increase in training data produced no measurable change. The v1 figure rests on only 37 wrong cases, which was never a firm enough baseline to compare against.
Two further cautions:
max_seq_length 2048→4096, eval engine Unsloth→vLLM, eval set 30+30→100+100),
so even a significant difference could not have been attributed to any one of
them.The reflect turn must carry the real verifier error, under the tool role,
using the same template as training — a mismatch between training and inference
prompting reintroduces the failure this adapter was built to remove.
Chen, K-Y., Su, F-Y., Chiang, J-H. (2026). The Self-Correction Illusion: LLMs Correct Others but Not Themselves. arXiv:2606.05976