Downloads · 30 days
100
19% of all-time downloads
sttempler/KORPatent-BGE
KORPatent-BGE is a sentence similarity model from sttempler. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as mit.
This model is a KOREAN patent-domain embedding model fine-tuned from BGE-M3 using curriculum triplet learning without relying on pseudo similarity scores (pseudo labels) generated by other models.
Downloads · 30 days
100
19% of all-time downloads
All-time downloads
517
Public
Parameters
568M
2.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
This model is a KOREAN patent-domain embedding model fine-tuned from BGE-M3 using curriculum triplet learning without relying on pseudo similarity scores (pseudo labels) generated by other models.
Instead of collecting dense, per-pair similarity scores for massive patent pairs—which is costly and often infeasible—we define Anchor / Positive / Negative triplets using objective rules based on:
Training progresses from easy discrimination to hard discrimination via 5-stage Curriculum Learning, encouraging the model to learn both coarse technical boundaries and fine-grained technical distinctions.
이 모델은 BGE-M3를 기반으로, 다른 모델이 생성한 의사 유사도 점수(의사 레이블, pseudo label)에 의존하지 않고 커리큘럼 트리플렛 학습(curriculum triplet learning)으로 미세조정한 한국어 특허 도메인 임베딩 모델입니다.
방대한 특허 쌍에 대해 쌍마다 조밀한 유사도 점수를 수집하는 방식은 비용이 크고 종종 실현이 불가능합니다. 이런 방식 대신, 다음의 객관적 규칙에 기반해 앵커(Anchor) / 포지티브(Positive) / 네거티브(Negative) 트리플렛을 정의했습니다.
Recommended
Not recommended / out of scope
권장
비권장 / 적용 범위 밖
Patent documents inherently share highly overlapping terminology, making it fundamentally difficult to separate them based on surface-level text alone.
As an illustrative example to demonstrate our model's capability to capture true semantic context, we selected five IPC subclasses for UMAP visualization: G01D (General Measurement), G01R (Measuring Electric Variables), H03K (Pulse Techniques), G05B (Control Systems), and H01L (Semiconductor Devices).
These subclasses represent a continuous technical workflow: Measurement → Signal Processing → Control → Hardware Manufacturing. While their textual vocabularies are heavily intertwined, they have distinctly different functional roles from a domain knowledge perspective.
특허 문서는 본질적으로 매우 중첩된 용어를 공유하기 때문에, 표면적인 텍스트만으로 구분하기가 근본적으로 어렵습니다. 모델이 실제 의미적 맥락을 포착하는 능력을 보여주기 위한 예시로, UMAP 시각화에 다섯 개의 IPC 서브클래스를 선택했습니다: G01D(일반 측정), G01R(전기 변량 측정), H03K(펄스 기술), G05B(제어 시스템), H01L(반도체 소자). 이 서브클래스들은 측정 → 신호 처리 → 제어 → 하드웨어 제조로 이어지는 연속적인 기술 흐름을 나타냅니다. 텍스트 어휘는 서로 크게 얽혀 있지만, 도메인 지식 관점에서는 기능적 역할이 뚜렷이 다릅니다.
from sentence_transformers import SentenceTransformer
import torch
model = SentenceTransformer("sttempler/KORPatent-BGE", device="cuda" if torch.cuda.is_available() else "cpu")
sentences = [
"인간 선호 데이터를 이용하여 상담 챗봇 응답을 정렬하는 RLHF 기반 학습 방법.",
"선호 피드백으로 보상 신호를 구성하여 생성 모델을 원하는 방향으로 학습시키는 방법."
]
emb = model.encode(sentences, normalize_embeddings=True)
score = float(emb[0] @ emb[1].T)
print(score)
# 0.7382504343986511
If you use this model in your research, please cite the following doctoral dissertation:
@phdthesis{kim2026koreanpatentembedding,
title = {A Study on Domain-Specific Long-Context Embedding Models for Advancing Korean Patent Retrieval},
author = {Kim, Yongwoo},
school = {Hanyang University},
department = {Department of Technology Management},
address = {Seoul, Korea},
year = {2026},
type = {Doctoral dissertation},
note = {Korean title: 한국어 특허 검색을 위한 장문 임베딩 모델},
annote = {https://huggingface.co/sttempler/KORPatent-BGE}
}