Downloads · 30 days
147
25% of all-time downloads
BurnyCoder/qwen2.5-coder-0.5b-swe-rl
qwen2.5-coder-0.5b-swe-rl is a text generation model from BurnyCoder. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains the merged output of the MiniBug-RL hidden-unit-test GRPO experiment. It targets single-function Python repair; it is not a repository-scale autonomous software-engineering agent.
Downloads · 30 days
147
25% of all-time downloads
All-time downloads
584
Public
Parameters
494M
1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1 GB · 98%
From the Hugging Face model README
This repository contains the merged output of the MiniBug-RL hidden-unit-test GRPO experiment. It targets single-function Python repair; it is not a repository-scale autonomous software-engineering agent.
The model, adapter, and measured results were originally published at immutable revision
5b6e22a4c6c01bec95d10e93a0fc78666eb9c543.
This later README-only correction fixes metric terminology and narrows claims to what the
recorded evidence establishes. It did not rerun training or evaluation. The audited
repository documentation is at GitHub commit
902ca72a46bbc55b3154a3e2f1cd5c8cff6ddde2.
Qwen/Qwen2.5-Coder-0.5B-Instruct at ea3f2471cf1b1f0db85067f1ef93848e38e88c25train chosen only on the 12-task validation split| Metric | Base | Selected model |
|---|---|---|
| Greedy pass@1 | 0.4167 | 0.5000 |
| Greedy hidden-test fraction | 0.6083 | 0.6042 |
| Observed sampled success@4 | 0.8333 | 0.8333 |
Observed sampled success@4 is the fraction of tasks with at least one complete repair among exactly four generated samples; it is not an unbiased pass@k estimator.
Paired hidden-test-fraction difference: -0.0042 with paired percentile-bootstrap 95% interval [-0.2917, 0.2625].
The validation learning gate, pre-specified in the producing source commit, passed. These small synthetic-split measurements should not be generalized to SWE-bench.
| Metric | Base | Selected model |
|---|---|---|
| Greedy pass@1 (164 Python repairs) | 0.2256 | 0.2317 |
| Timeouts | 3 | 4 |
Protocol: bigcode/humanevalpack at immutable dataset revision 9a41762f73a8cb23bb5811b73d5aab164efcf378, prompt variant humanevalfixdocs-python, greedy n=1.
Postprocessing reference: pinned BigCode Python harness.
Paired pass@1 difference: 0.0061 with paired percentile-bootstrap 95% interval [-0.0183, 0.0366].
HumanEvalPack examples, candidate outcomes, and scores did not enter the implemented training, reward-computation, checkpoint-selection, or tuning data flows. Its pinned harness source defined the frozen external protocol; public benchmark contamination may still affect both policies.
Executable candidates ran under MiniBug-RL's host-enforced 3-second wall-clock deadline around resource-limited isolated Docker, rather than the pinned BigCode Python harness's 10-second limit; treat this as a MiniBug-sandbox measurement, not a directly comparable leaderboard score.
Sandbox image: sha256:a869cd1dffb8c87afad1bb1302106cb9f5cb580641c7391bb73f4ab077f140d9.
Docker controls are defense in depth, not proof that this runner is safe for arbitrary hostile code. Internal parser/policy rejections and deterministic reward cache hits do not start containers.
Paired percentile-bootstrap intervals describe these fixed task samples; they do not by themselves establish broad model quality.
results.json: candidate-level outcomes and aggregates. Its immutable historical key sampled_pass_at_4 denotes the observed four-draw statistic relabeled above.training_config.json: resolved non-secret configuration.from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BurnyCoder/qwen2.5-coder-0.5b-swe-rl"
result_revision = "5b6e22a4c6c01bec95d10e93a0fc78666eb9c543"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=result_revision)
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision=result_revision,
dtype="auto",
)
The separately loadable LoRA adapter and tokenizer are in adapter/.
Candidate-level measurements and the resolved configuration are included.