Downloads · 30 days
21
16% of all-time downloads
jjp97/laal-embedding-v0
laal-embedding-v0 is a sentence similarity model from jjp97. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
laal-embedding-v0 is a Sentence-Transformers embedding model fine-tuned from intfloat/multilingual-e5-large-instruct for improved retrieval-oriented semantic search, with a focus on Korean fire-safety and legal-domain…
Downloads · 30 days
21
16% of all-time downloads
All-time downloads
131
Public
Parameters
560M
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.2 GB · 99%
From the Hugging Face model README
laal-embedding-v0 is a Sentence-Transformers embedding model fine-tuned from
intfloat/multilingual-e5-large-instruct for improved retrieval-oriented semantic search, with a focus on Korean fire-safety and legal-domain text.
intfloat/multilingual-e5-large-instruct⚠️ Important This model uses fixed instruction prefixes defined in
config_sentence_transformers.json. Always pass raw text toencode(). Do NOT manually prepend instruction strings.
This model applies different fixed prefixes depending on the input type.
Instruct: Given a web search query, retrieve relevant passages that answer the query.
Query:
title: none
text:
These prefixes are automatically applied by Sentence-Transformers via
config_sentence_transformers.json.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("jjp97/laal-embedding-v1")
q_emb = model.encode_query("화재 시 대피 방법")
p_emb = model.encode_document("화재가 발생하면 즉시 119에 신고하고 안전한 경로로 대피해야 한다.")
# Do NOT do this
q = "Instruct: Given a web search query, retrieve relevant passages...\nQuery: 화재 시 대피 방법"
emb = model.encode(q)
Contrastive learning (InfoNCE)
In-batch negatives
Temperature (tau): 0.05
Regularization: GOR (spread-out loss)
gor_lambda = 0.001gor_max_samples = 64Training examples: 43,983
Format: (query, positive passage)
Hard negatives: enabled
max_hn_per_example_train = 2Training data consists of domain-specific Korean fire-safety and legal documents (private / curated dataset).
This model follows the standard Sentence-Transformers pipeline:
Query–passage cosine similarity shows reasonable separation between relevant and irrelevant passages.
This model is intended for evaluation on the MTEB leaderboard. When reporting results, please specify:
jjp97/laal-embedding-v0If you use this model, please cite:
@misc{laal_embedding_v0_2025,
title = {laal-embedding-v0},
author = {Park, Jeongjae},
year = {2025},
howpublished = {Hugging Face model card},
}