Downloads · 30 days
5
36% of all-time downloads
BeastxD/text2cypher_lora_v6_shaped
text2cypher_lora_v6_shaped is a reinforcement learning model from BeastxD. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Overnight experiment, 2026-08-25/26: GRPO on top of BeastxD/text2cypherlorav7, using a shaped, partial-credit reward instead of the binary +1/0/-0.5 reward the first GRPO attempt (BeastxD/text2cypherlorav6grpo) used.
Downloads · 30 days
5
36% of all-time downloads
All-time downloads
14
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
Overnight experiment, 2026-08-25/26: GRPO on top of BeastxD/text2cypher_lora_v7, using a
shaped, partial-credit reward instead of the binary +1/0/-0.5 reward the first GRPO
attempt (BeastxD/text2cypher_lora_v6_grpo) used.
| comparison | execution accuracy | delta | McNemar p | verdict |
|---|---|---|---|---|
| v7 → v6_shaped, Neo4j 2024v1 | 55.66% → 56.84% | +1.18pp | 0.0890 | not significant (p<0.05) |
| v6_grpo (binary) → v6_shaped | 56.21% → 56.78% | +0.57pp | 0.3100 | not significant |
Do not read this as "shaped reward beats v7" — p=0.089 does not clear the 0.05 bar this project uses everywhere else. What is worth reading into it: the original binary-reward run scored p=0.42 against v7 after 1,201 steps. This run reached p=0.089 after only 782 steps (65% as much training, from an earlier checkpoint-600, not even its own final state). That is a materially stronger signal on less compute, which is evidence in favor of running the shaped reward longer -- not proof it already works.
The binary reward gives an identical 0.0 to every wrong-but-executing rollout, regardless
of how close it is. GRPO only learns from within-group variance, so two rollouts that are
both wrong in different ways score the same and produce zero gradient. Measured on the
original run: usable-groups-per-step decayed 1.35 → 0.85 over its lifetime, and held-out
accuracy never moved (p=0.42).
+1.0 exact result-set match (unchanged)
0.8×J wrong, J = Jaccard overlap between predicted and gold result multisets
-0.5 execution error (unchanged)
See v6/code/grpo_reward_shaped.py for the implementation and validation.
Fresh from v7 (not resumed from the binary run, to isolate reward as the only variable).
Self-limiting deadline (timezone for when results were needed was unknown), stopped after
782 steps / ~2h of training -- short of the 5,311-step epoch and short of the original
run's 1,201 steps. frac_reward_zero_std still rose over the run (0.320 → 0.480), similar
in shape to the binary run, so the shaping did not eliminate the flat-group problem --
it appears to have made the non-flat groups more informative instead.
This is a partial run merging an early checkpoint (step 600 of 782 reached). A longer, properly-scheduled continuation is the natural next step, not a verdict on the method.
Same prompt/pruning contract as v7 and v6_grpo -- v7's ~618-char system prompt with exact-match schema pruning.