Downloads · 30 days
0
onesfour/MiniOneRec-1.5B-SFT-GDPO
MiniOneRec-1.5B-SFT-GDPO is a text generation model from onesfour. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Generative Recommendation Model — fine-tuned from Qwen2.5-1.5B-Instruct on Amazon Industrial & Scientific reviews.
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
12.4 GB
Likes
0
Public
Click a slice to open those files.
.safetensors6.6 GB · 54%
From the Hugging Face model README
Generative Recommendation Model — fine-tuned from Qwen2.5-1.5B-Instruct on Amazon Industrial & Scientific reviews.
📖 GitHub: YuyaoFan/MiniOneRec-1.5B
| Directory | Description | Usage |
|---|---|---|
sft_model/ | RQ-VAE + SFT model weights | AutoModelForCausalLM.from_pretrained() |
gdpo_best_checkpoint-1155/ | Best GDPO RL checkpoint (step 1155) | Best RL model for evaluation |
data/ | Dataset files (.npy, .npz, .csv, .inter) | Required for training reproduction |
This model transforms item recommendation into a sequence generation task using Semantic IDs (SIDs):
User History → Qwen2.5-1.5B → Semantic ID Sequence → Item Recommendation
Evaluation on Amazon Industrial & Scientific test set (beam=50):
| HR@3 | NDCG@3 | HR@5 | NDCG@5 | HR@10 | NDCG@10 |
|---|---|---|---|---|---|
| 0.0904 | 0.0792 | 0.1061 | 0.0856 | 0.1337 | 0.0945 |
| Method | Step | HR@5 | NDCG@5 | HR@10 | NDCG@10 | vs SFT |
|---|---|---|---|---|---|---|
| GRPO | 1320 | 0.1136 | 0.0926 | 0.1379 | 0.1005 | +6.4% |
| GDPO | 1155 | 0.1050 | 0.0880 | 0.1253 | 0.0944 | -0.1% |
GRPO significantly outperforms SFT (+6.4% NDCG@10). GDPO shows limited gains in this sparse-reward recommendation setting.
| Method | HR@5 | NDCG@10 |
|---|---|---|
| SASRec | 0.0909 | 0.0806 |
| TIGER | 0.1010 | 0.0908 |
| D3 | 0.1213 | 0.1082 |
| MiniOneRec (7B, paper) | 0.1321 | 0.1167 |
| Ours (1.5B, GRPO) | 0.1136 | 0.1005 |
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"onesfour/MiniOneRec-1.5B-SFT-GDPO",
subfolder="sft_model", # or "gdpo_best_checkpoint-1155"
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
"onesfour/MiniOneRec-1.5B-SFT-GDPO",
subfolder="sft_model",
trust_remote_code=True,
)
prompt = """Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.
### Instruction:
Can you predict the next possible item that the user may expect?
### User Input:
The user has interacted with items <a_115><b_58><c_12>, <a_86><b_17><c_25> in chronological order. Can you predict the next possible item that the user may expect?
### Response:
"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=10, num_beams=50)
predicted = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(f"Predicted SID: {predicted}")
Items are represented as 3-level Semantic IDs with an optional deduplication suffix:
<a_X>: Level 0 code (0-255)<b_Y>: Level 1 code (0-255)<c_Z>: Level 2 code (0-255)<d_W>: Deduplication suffix (if collision exists)Example: <a_115><b_58><c_12> represents a unique item in the catalog.
To map predicted SIDs back to item titles, use data/Amazon/index/Industrial_and_Scientific.item.json and data/Amazon/info/Industrial_and_Scientific_5_2016-10-2018-11.txt from the repository.
transformers>=4.57.1
torch>=2.6.0
See YuyaoFan/MiniOneRec-1.5B for full source code, scripts, and documentation.
Quick start:
git clone https://github.com/YuyaoFan/MiniOneRec-1.5B.git
cd MiniOneRec-1.5B
# Download data files from this HF repo → data/Amazon/
# Download model from this HF repo/sft_model → output/sft/.../final_checkpoint/
conda create -n MiniOneRec python=3.11 -y && conda activate MiniOneRec
bash scripts/install_deps.sh
bash scripts/run_rqvae_sft_grpo.sh
@misc{kong2025minionerec,
title={MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation},
author={Xiaoyu Kong and Leheng Sheng and Junfei Tan and Yuxin Chen and
Jiancan Wu and An Zhang and Xiang Wang and Xiangnan He},
year={2025},
eprint={2510.24431},
archivePrefix={arXiv},
primaryClass={cs.IR},
}
@misc{liu2026gdpo,
title={GDPO: Group reward-Decoupled Normalization Policy Optimization},
author={Shih-Yang Liu et al.},
year={2026},
eprint={2601.05242},
archivePrefix={arXiv},
}
This model inherits Apache 2.0 from Qwen2.5-1.5B-Instruct.