Downloads · 30 days
12
50% of all-time downloads
Guardrium/spicy-motivator-ppo
spicy-motivator-ppo is a reinforcement learning model from Guardrium. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
한국어 명언을 비꼬는 문장으로 변환하는 모델 (PPO/REINFORCE로 학습)
Downloads · 30 days
12
50% of all-time downloads
All-time downloads
24
Public
Repo size
185 MB
Likes
0
Public
Click a slice to open those files.
.safetensors168 MB · 91%
From the Hugging Face model README
한국어 명언을 비꼬는 문장으로 변환하는 모델 (PPO/REINFORCE로 학습)
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
# Base 모델 로드
base_model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
torch_dtype=torch.float16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
# LoRA 어댑터 로드
model = PeftModel.from_pretrained(base_model, "YOUR_USERNAME/spicy-motivator-ppo")
# 생성
prompt = "### 명언: 노력은 배신하지 않는다.\n### 비꼬는 답변:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))