Downloads · 30 days
7
39% of all-time downloads
elevateecho/sn120-9fbb78acd964
sn120-9fbb78acd964 is a image-text-to-text model from elevateecho. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Reason-v4 GRPO checkpoint trained in the /mining/ralph loop on an 8×H100 PCIe box. Pushed for a live SN120 duel against the sitting king vera6/affine-5g4yy75zuz-t6 @ 8e3f1695e058837ed80fec3238ff439fdc2d0f0e (reign 36)…
Downloads · 30 days
7
39% of all-time downloads
All-time downloads
18
Public
Parameters
34.7B
70.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors70.2 GB · 100%
From the Hugging Face model README
Reason-v4 GRPO checkpoint trained in the /mining/ralph loop on an 8×H100
PCIe box. Pushed for a live SN120 duel against the sitting king
vera6/affine-5g4yy75zuz-t6 @ 8e3f1695e058837ed80fec3238ff439fdc2d0f0e
(reign 36). Architecture is stock Qwen3_5MoeForConditionalGeneration
(hidden 2048, 40 layers, 256 experts / 8 active). No custom modeling code,
no auto_map, no *.py.
vera6/affine-5g4yy75zuz-t6 @ 8e3f1695e058837ed80fec3238ff439fdc2d0f0ecand_p24-grpo-king-affine-5g4yy75zu).
Local v4 screen vs t6: n=400, margin −0.00009, z −0.09, med |z| 139,
B-pass 0.46 (tie / BELOW_BAR)./mining/sim/merge_lora2.py (nonzero delta + 333 visual tensors).Experiment path: /mining/ralph/runs/p28-grpo-p24-grpo-king-affine-5g4
Teacher-anchored Reason v4 GRPO (train_reason_grpo.py).
Per-sample reward:
a_i = lpC(y_i | z_A) − lpC(y_i | ∅) # k=3 teacher refs
Reason = τ · log((1/k) · Σ exp(a_i/τ)) # τ=0.03
reward = Reason + length_shape(|z|) # penalize |z|≥220 only
Winner-only tail-boost 2.0 on the best group member. Ranked quantity is the
thought z (action y is not the score). Teacher is the frozen Affine
teacher zai-org/GLM-4.5-Air-FP8, two local vLLM TP=2 endpoints.
/mining/ralph/data/grpo.jsonlTHOUGHT + last closed bash fence)| knob | value |
|---|---|
| method | HiAlpha-GRPO (LoRA) |
| lr | 5e-6 |
| LoRA r / α / dropout | 16 / 128 / 0.05 |
| target modules | q,k,v,o,gate,up,down _proj |
| group size G | 4 |
| steps | 200 |
| max new tokens (z sample) | 64 |
| max seq | 6144 |
| KL coef | 0.0 |
| tail-boost | 2.0 winner-only |
| length shape | penalty on |z|≥220, no length bonus |
| τ / k_refs | 0.03 / 3 |
| dtype | bfloat16 |
Train wall: 4524 s (~75 min). Last-20 mean reward 0.044. Trainable 8.36M /
34.7B (0.024%). GPUs 4–7 for LoRA (device_map=auto); teachers on 0–1 and 2–3.
Screened 2026-08-19 against the same king SHA still sitting at push time
(vera6/affine-5g4yy75zuz-t6 @ 8e3f1695…). gate_screen.py / fast_screen
n=160:
| cand | king | |
|---|---|---|
| mean Reason | 0.01184 | 0.00995 |
| med |z| | 144 | 145 |
| B-pass | 0.46 | 0.49 |
p28 was the closest this-loop candidate vs t6. Live duel is n=1300 k=3; this n=160 slice does not prove a crown.
Same family the eval pod loads: Qwen3.5-MoE, canonical sharded safetensors
(model-00001-of-00002 + model-00002-of-00002 + model.safetensors.index.json
model-visual.safetensors). No --trust-remote-code. Local screen served
this merge with vLLM TP=2.8× NVIDIA H100 80GB PCIe. Merge on CPU (device_map=cpu).