Downloads · 30 days
21
2% of all-time downloads
BytedTsinghua-SIA/JustRL-R1-7B
JustRL-R1-7B is a text generation model from BytedTsinghua-SIA. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
JustRL-R1-7B is a research checkpoint from the Direct-OPD collection. It is released for reproducible research on post-training and reinforcement-learning-style optimization for reasoning-oriented language models.
Downloads · 30 days
21
2% of all-time downloads
All-time downloads
859
Public
Parameters
7.6B
15.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.2 GB · 100%
From the Hugging Face model README
JustRL-R1-7B is a research checkpoint from the Direct-OPD collection. It is released for reproducible research on post-training and reinforcement-learning-style optimization for reasoning-oriented language models.
BytedTsinghua-SIA/JustRL-R1-7Bc7wc7w/20260603_JustRL_R1_7B_AdaptiveKL_EPS0p01_LEN2k_ckpt280A JustRL Direct-OPD checkpoint for an R1 7B model family, using the adaptive-KL setting indicated by the source checkpoint name.
The checkpoint was mirrored from ModelScope to Hugging Face for easier access and inclusion in the Direct-OPD collection. The source checkpoint name records the available release metadata, including method, model family, KL/adaptive-KL setting when present, sequence length setting, and checkpoint step. Full training data, hyperparameters, and evaluation protocol should be taken from the associated Direct-OPD release materials when available.
This model is intended for research use, including:
It is not intended for direct deployment in safety-critical, medical, legal, financial, or other high-stakes settings without independent evaluation and safeguards.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BytedTsinghua-SIA/JustRL-R1-7B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "user", "content": "Explain the main idea of Direct-OPD in one paragraph."},
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
output_ids = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
top_p=0.95,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
No benchmark results are included in this model card yet. Users should evaluate the checkpoint on the tasks and safety criteria relevant to their use case before drawing conclusions or deploying derived systems.
Recommended reporting for downstream evaluations:
Citation information for the associated Direct-OPD research release will be added when available. If you use this checkpoint, please cite the Direct-OPD project or paper once published, and include the model repository URL in your reproducibility artifacts.
This checkpoint is part of the Direct-OPD model collection maintained by BytedTsinghua-SIA and mirrored from the original ModelScope release.