Downloads · 30 days
13
52% of all-time downloads
djByun/TPGTP
TPGTP is a machine learning model from djByun. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
이 모델은 논문 제목을 입력하면 해당 논문이 발표될 가능성이 높은 학술대회를 예측하는 한국어 경량 LLM입니다. Agent AI 활용 확산과 맞물려, 연구현장에서 자연어 기반의 분류 업무를 자동화할 수 있도록 실무 데이터를 기반으로 구축하였습니다.
Downloads · 30 days
13
52% of all-time downloads
All-time downloads
25
Public
Repo size
143 MB
Likes
0
Public
Click a slice to open those files.
.json104 MB · 51%
From the Hugging Face model README
이 모델은 논문 제목을 입력하면 해당 논문이 발표될 가능성이 높은 학술대회를 예측하는 한국어 경량 LLM입니다.
Agent AI 활용 확산과 맞물려, 연구현장에서 자연어 기반의 분류 업무를 자동화할 수 있도록 실무 데이터를 기반으로 구축하였습니다.
본 프로젝트는 정보통신기획평가원(IITP)의 정책 수혜자로서, 실제 기관에서 직면한 '논문-학술대회 분류' 업무를 효율화하는 데 기여하고자 기획되었습니다.
google/gemma-3-1b-it한국연구재단_학술대회논문심사_20241231.csv{"text": 논문 제목, "label": 학술대회명} 형태의 JSONL 변환[INST] 논문 제목: {제목} 어떤 학술대회명인가요? [/INST] {학술대회명} 형식으로 Prompt 생성from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained("JeongHeum/gemma3-korean-academic-classifier")
tokenizer = AutoTokenizer.from_pretrained("JeongHeum/gemma3-korean-academic-classifier")
prompt = "[INST] 논문 제목: 딥러닝 기반 한국어 음성 인식 시스템 [/INST]"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=20)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
# 예시 출력: 한국음성처리학회