Downloads · 30 days
0
ZhengyanWan/dFlowGRPO-ScienceQA
dFlowGRPO-ScienceQA is a machine learning model from ZhengyanWan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
LoRA adapter for FUDOKI, fine-tuned with dFlowGRPO on the ScienceQA reward. Training step: 1500.
Downloads · 30 days
0
Access
Public
Updated May 8, 2026
Repo size
182 MB
Likes
0
Public
Click a slice to open those files.
.pt90.9 MB · 50%
From the Hugging Face model README
LoRA adapter for FUDOKI,
fine-tuned with dFlowGRPO on the ScienceQA reward.
Training step: 1500.
See WanZhengyan/dFlowGRPO for training & evaluation code.
lora_adapterema_state.pthuggingface-cli download ZhengyanWan/dFlowGRPO-ScienceQA \
--local-dir output_grpo_u_scienceqa_kl0.01/checkpoint_0_step1500
Then run the matching evaluator from discrete_flow_grpo/.