Downloads · 30 days
18
14% of all-time downloads
THU-KEG/LLaDA-8B-BGPO-code
LLaDA-8B-BGPO-code is a reinforcement learning model from THU-KEG. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
[](https://arxiv.org/abs/2510.11683) [](https://github.com/THU-KEG/BGPO)
Downloads · 30 days
18
14% of all-time downloads
All-time downloads
131
Public
Parameters
8B
16 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16 GB · 100%
From the Hugging Face model README
LLaDA-8B-BGPO-code is an 8-billion parameter diffusion large language model (dLLM) that was trained on LLaDA-8B-Instruct using Boundary-Guided Policy Optimization (BGPO) for enhanced code generation capabilities.