Downloads · 30 days
531
100% of all-time downloads
ranjanrajib/DPO-trained-Qwen2.5-Python-Coder-7B
DPO-trained-Qwen2.5-Python-Coder-7B is a text generation model from ranjanrajib. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This model is a Python-focused, DPO fine-tuned version of Qwen2.5-Coder-7B-Instruct, designed to improve the model's ability to generate functionally correct Python code.
Downloads · 30 days
531
100% of all-time downloads
All-time downloads
531
Public
Parameters
7.6B
15.6 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors15.6 GB · 100%
From the Hugging Face model README
This model is a Python-focused, DPO fine-tuned version of Qwen2.5-Coder-7B-Instruct, designed to improve the model's ability to generate functionally correct Python code.
The model was trained using Direct Preference Optimization (DPO) with automatically constructed preference pairs. Candidate Python solutions were generated for programming problems and evaluated by executing them against unit tests. Solutions that passed the required tests were used as preferred examples, while solutions that failed tests were used as rejected examples.
The resulting model was evaluated against the original Qwen2.5-Coder-7B-Instruct model on held-out programming problems.
The base model:
Qwen/Qwen2.5-Coder-7B-Instruct
was fine-tuned using: DPO + LoRA + execution-based preference data
The resulting model demonstrates improved coding performance on the evaluated held-out benchmark:
| Metric | Qwen2.5-Coder-7B-Instruct | DPO Model | Improvement |
|---|---|---|---|
| Pass@1 | 41.87% | 45.22% | +3.35 pp |
| Pass@5 | 48.87% | 51.10% | +2.23 pp |
| Pass@10 | 50.99% | 53.16% | +2.17 pp |
| Tests passed per answer | 66.20% | 69.41% | +3.21 pp |
The primary result is the improvement in Pass@1, meaning that a single generated solution was more likely to produce a correct solution on the evaluated programming problems. The model also showed improved Pass@5, Pass@10, and partial test coverage.
The base model and the merged trained model use the same VRAM. The merge changes weight values, not tensor shapes: both models have the same 7.6B parameters and architecture, so the model weights occupy the same amount of bf16 memory.
| Component | bf16 VRAM |
|---|---|
| Weights (base or merged D4) | 14.19 GiB |
| KV cache, per token | 56 KiB |
| Activations / workspace | ~0.5–1 GiB |
| Practical single-stream total | ~15–16 GiB |
Base model: Qwen/Qwen2.5-Coder-7B-Instruct at snapshot c03e6d358207e414f1eca0bb1891e29f1db0e242. The merge was performed against exactly these weights.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"ranjanrajib/DPO-trained-Qwen2.5-Python-Coder-7B", dtype="bfloat16", device_map="auto")
tok = AutoTokenizer.from_pretrained("ranjanrajib/DPO-trained-Qwen2.5-Python-Coder-7B")
msgs = [{"role": "user", "content": "Write a Python function that merges overlapping intervals."}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(text, return_tensors="pt").to(model.device), max_new_tokens=512)
print(tok.decode(out[0], skip_special_tokens=True))
Or fetch the files without loading them (~14 GB, four shards):
hf download ranjanrajib/DPO-trained-Qwen2.5-Coder-7B --exclude "adapter/*" --local-dir ./d4
The adapter lives in adapter/. Applying it to the base model gives the same result as the merged weights, at a fraction of the download size.
hf download ranjanrajib/DPO-trained-Qwen2.5-Python-Coder-7B --include "adapter/*" --local-dir ./d4-lora
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM
BASE = "Qwen/Qwen2.5-Coder-7B-Instruct"
REVISION = "c03e6d358207e414f1eca0bb1891e29f1db0e242" # the snapshot this was trained against
base = AutoModelForCausalLM.from_pretrained(
BASE, revision=REVISION, dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(
base, "DPO-trained-Qwen2.5-Python-Coder-7B", subfolder="adapter"
)
The prompt
PROMPT = """You are an expert Python programmer.
Solve the following programming problem.
Problem:
Calculate factorials for a list of numbers in parallel using multiprocessing.
It should return:
dict[int, int]: A dictionary with numbers as keys and their factorial as values.
It should raise:
ValueError: If any element in the input list is not an integer or is negative.
It may use:
multiprocessing.Pool
math.factorial.
Write a function with this exact signature:
def task_func(numbers: list) -> dict:
It must satisfy:
>>> factorials = task_func([5, 6, 7, 8, 9])
>>> factorials[5] == 120 and factorials[9] == 362880
True
Required function signature:
def task_func(numbers: list) -> dict:
Requirements:
- Implement the requested function.
- Follow the function signature.
- Handle the specified edge cases.
- Use Python.
- Return only the implementation.
- Do not provide an explanation.
- Do not use eval().
- Do not use exec().
- Do not perform network operations.
- Do not read or write files.
"""
### Generating
import re
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, set_seed
MODEL = "ranjanrajib/DPO-trained-Qwen2.5-Coder-7B"
model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained(MODEL)
text = tok.apply_chat_template(
[{"role": "user", "content": PROMPT}], tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
set_seed(159000) # the seed this particular sample was drawn with
out = model.generate(
**inputs,
do_sample=True,
temperature=0.2,
top_p=0.95,
max_new_tokens=512,
pad_token_id=tok.eos_token_id,
)
response = tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
code = re.search(r"```python\n(.*?)```", response, re.S).group(1)
print(code)
### What the model returned
84 generated tokens, reproduced verbatim:
import multiprocessing
import math
def factorial_task(number):
if not isinstance(number, int) or number < 0:
raise ValueError(f"Invalid input: {number}")
return number, math.factorial(number)
def task_func(numbers: list) -> dict:
with multiprocessing.Pool() as pool:
results = pool.map(factorial_task, numbers)
return dict(results)
If your primary use case is Python code generation, this model is intended to be a stronger alternative to the original Qwen2.5-Coder-7B-Instruct model based on the evaluation performed in this project.
It is particularly relevant for users who want to:
The main motivation for the fine-tuning was not simply to make the model generate more code, but to encourage the model toward solutions that work. The preference data therefore used actual program execution and unit-test results rather than subjective judgments about whether generated code looked correct.
The underlying model architecture and capabilities come from Qwen2.5-Coder-7B-Instruct. This model adds a preference-optimization stage on top of that model.
Conceptually:
Qwen2.5-Coder-7B-Instruct
│
▼
Generate multiple Python
candidate solutions
│
▼
Execute candidates
against tests
│
┌────┴────┐
▼ ▼
Pass Fail
│ │
▼ ▼
Chosen Rejected
│ │
└────┬────┘
▼
DPO
│
▼
LoRA Adapter
│
▼
Python-focused model
The goal is to shift the model toward coding behaviors associated with solutions that successfully execute against the available tests.
This is not a new foundation model. It is a parameter-efficient fine-tuned adapter for Qwen2.5-Coder-7B-Instruct. The base model remains unchanged.
Qwen/Qwen2.5-Coder-7B-InstructThe model was evaluated against the original Qwen2.5-Coder-7B-Instruct using a controlled benchmark consisting of 508 held-out BigCodeBench programming problems. The same prompts, generation settings, random seeds, number of samples, and grading procedure were used for both models.
| Metric | Base Model | DPO Model |
|---|---|---|
| Pass@1 | 41.87% | 45.22% |
| Pass@5 | 48.87% | 51.10% |
| Pass@10 | 50.99% | 53.16% |
The improvement in Pass@1 is approximately 3.35 percentage points.
The average proportion of tests passed by each generated answer increased from 66.20% → 69.41%. This is a secondary metric intended to provide additional information about functional correctness beyond binary problem-level pass/fail.
The evaluation used the same 508 problems for both models, allowing direct comparison of model behavior on individual problems. For the first generated answer:
| Base → DPO | Problems |
|---|---|
| Correct → Correct | 186 |
| Correct → Incorrect | 17 |
| Incorrect → Correct | 41 |
| Incorrect → Incorrect | 264 |
The DPO model converted 41 problems from incorrect to correct while 17 previously correct problems became incorrect. This provides additional evidence that the aggregate improvement is associated with meaningful changes in problem-level behavior rather than simply differences in the evaluated problem sets.
Developers interested in experimenting with a locally hosted code-generation model that has been specifically fine-tuned toward execution-validated solutions.
Researchers studying DPO, preference learning, LLM alignment, code-generation models, automated preference construction, execution-based evaluation, and parameter-efficient fine-tuning.
The model can also be useful as an experimental model for comparing base versus preference-tuned models, coding correctness, Pass@k, partial test coverage, and preference-learning behavior.
This model should not automatically be considered better for every possible task. The experiment specifically focused on Python code generation. If your application requires broad general-purpose instruction following, other programming languages, or capabilities that were not evaluated here, you should benchmark both models for your specific workload.
Improved performance on the evaluated Python coding benchmark, not a guarantee of improvement across every coding or general-purpose task.
The model starts from Qwen/Qwen2.5-Coder-7B-Instruct containing approximately 7.7 billion parameters. Rather than updating all parameters, the experiment uses LoRA-based parameter-efficient fine-tuning. The resulting adapter contains approximately 80.7 million trainable parameters and is ~308 MB in size.
Direct Preference Optimization provides a method for optimizing a language model directly from preference pairs without requiring a separate reward-model training stage. For this experiment, a preference pair has the form:
Problem
│
├── Candidate A → passes tests → chosen
│
└── Candidate B → fails tests → rejected
DPO then learns to increase the relative likelihood of the chosen response compared with the rejected response.
The preference dataset was generated specifically for this experiment, constructed from programming problems drawn from BigCodeBench, MBPP, HumanEval, and handwritten programming problems.
| Dataset | License |
|---|---|
| BigCodeBench | Apache 2.0 |
| MBPP | CC BY 4.0 |
| HumanEval | MIT |
Candidate solutions were generated automatically using five strategies: normal, straightforward, edge-case-focused, alternative, and optimized. Every candidate was then executed in a Docker sandbox against the associated unit tests to form chosen/rejected preference pairs.
After filtering and deduplication:
The dataset was split by problem, rather than by individual preference pair, to prevent different candidate solutions from the same underlying programming problem from being distributed across training and validation. No human preference labels or external LLM judges were used; the preference signal was generated entirely from objective unit-test execution.
| Parameter | Value |
|---|---|
| Rank | 32 |
| Alpha | 64 |
| Dropout | 0.05 |
| Target modules | Attention and MLP projections |
| Epochs | 2 |
| Maximum sequence length | 1,024 |
| Batch size | 4 |
| Gradient accumulation | 4 |
| Effective batch size | 16 |
| Optimizer | AdamW |
| Gradient clipping | 1.0 |
| Gradient checkpointing | Enabled |
The primary evaluation benchmark consisted of 508 BigCodeBench programming problems reserved separately from the preference-training pool.
The DPO training objective produced a strong preference signal on the training examples, but validation results were substantially weaker:
| Metric | Training | Validation |
|---|---|---|
| Preference accuracy | ~84.4% | ~45.8% |
| Reward margin | Positive | Approximately neutral |
This indicates that the explicit preference signal learned during training did not generalize strongly to the validation preference examples, even though downstream coding performance improved on the held-out benchmark.
The results suggest that optimizing an automatically constructed preference objective can improve downstream coding performance even when explicit preference metrics show weak validation generalization. Potential explanations include:
This model was developed as an independent, out-of-office AI/ML research project. The complete experimental lifecycle:
Research Question
↓
Experimental Design
↓
Preference Data Generation
↓
Automated Evaluation
↓
DPO + LoRA Training
↓
Controlled Benchmarking
↓
Statistical / Paired Analysis
↓
Model Release
The trained LoRA adapter and technical documentation are available on Hugging Face: 🔗 ranjanrajib/DPO-trained-Qwen2.5-Coder-7B
Qwen/Qwen2.5-Coder-7B-Instruct. This is a modified derivative of that model by the Qwen team at Alibaba Cloud.