Downloads · 30 days
8
20% of all-time downloads
AGENTDARS/Reviewer-7B
Reviewer-7B is a machine learning model from AGENTDARS. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
Reviewer-7B is a fine-tuned on DeepSeek-R1-Distill-Qwen-7B, optimized for selecting the best patch among multiple patches generated by our DARS agent while solving software engineering problems.
Downloads · 30 days
8
20% of all-time downloads
All-time downloads
40
Public
Repo size
36.6 MB
Likes
1
Public
Click a slice to open those files.
.pt31.4 MB · 72%
From the Hugging Face model README
Reviewer-7B is a fine-tuned on DeepSeek-R1-Distill-Qwen-7B, optimized for selecting the best patch among multiple patches generated by our DARS agent while solving software engineering problems.
We use vLLM to deploy and infer the model. Please follow this tutorial here to use our LoRA weights with vLLM.
We use our code review dataset where each instance contains several git patches with critiques for each each patch. The model learns to generate critiques for multiple patches and select the best patch.
| Hyperparameter | Value |
|---|---|
| Training regime | BF16 mixed precision |
| Optimizer | AdamW with cosine learning rate scheduler |
| LoRA Configuration | rank=8, alpha=32, dropout=0.1 |
| Batch Size | 48 |
| Learning Rate | 1e-5 |
| Sequence Length | 14K tokens |
| Fine-tuning Epochs | 1 |
| Compute Environment | DeepSpeed for memory-efficient distributed training |
| Compute Infrastructure | 8x H100 |
We use training script provided in Qwen-2.5 codebase.
Using this model as a reviewer with DARS trajectories generated using Claude 3.5 Sonnet V2 achieves 38.7% on SWE-Bench Lite.