Downloads · 30 days
187
100% of all-time downloads
NGA-KR/ko-embed-v0
ko-embed-v0 is a sentence similarity model from NGA-KR. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
Korean dense text embedding model (149M parameters), fine-tuned from skt/A.X-Encoder-base.
Downloads · 30 days
187
100% of all-time downloads
All-time downloads
187
Public
Parameters
149M
297 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors297 MB · 100%
From the Hugging Face model README
Korean dense text embedding model (149M parameters), fine-tuned from skt/A.X-Encoder-base.
This model was trained only on public English retrieval datasets — no Korean benchmark train splits were used:
Consequently, the model is 100% zero-shot on MTEB(kor, v1) and MTEB(kor, v2) (declared via training_datasets in the MTEB model metadata).
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("NGA-KR/ko-embed-v0")
q = model.encode(["전세 계약서에 기간을 안 썼으면 얼마 동안 유효해?"], prompt_name="query")
d = model.encode(["임대차 계약 기간을 정하지 않은 경우 2년으로 본다..."], prompt_name="document")
print(q @ d.T)
Prompts are stored in the model config: query: for queries, passage: for documents.
Evaluated with mteb on MTEB(kor, v2) (20 tasks). See the MTEB leaderboard entry for full per-task scores.
apache-2.0. Base model license: see skt/A.X-Encoder-base.