Downloads · 30 days
650
4% of all-time downloads
SamilPwC-AXNode-GenAI/PwC-Embedding_expr
PwC-Embedding_expr is a sentence similarity model from SamilPwC-AXNode-GenAI. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
We trained the PwC-Embedding-expr model on top of the multilingual-e5-large-instruct embedding model. To enhance performance in Korean, we applied our curated augmentation to STS datasets and fine-tuned the E5 model u…
Downloads · 30 days
650
4% of all-time downloads
All-time downloads
15.5K
Public
Parameters
560M
6.7 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors2.2 GB · 99%
From the Hugging Face model README
We trained the PwC-Embedding-expr model on top of the multilingual-e5-large-instruct embedding model.
To enhance performance in Korean, we applied our curated augmentation to STS datasets and fine-tuned the E5 model using a carefully balanced ratio across datasets.
⚠️ This is an experimental model and is under continuous development.
PwC-Embedding_expr was evaluated on the Korean subset of MTEB.
A leaderboard link will be added once it is published.
| Task | PwC-Embedding_expr |
|---|---|
| KLUE-STS | 0.88 |
| KLUE-TC | 0.73 |
| Ko-StrategyQA | 0.80 |
| KorSTS | 0.84 |
| MIRACL-Reranking | 0.72 |
| MIRACL-Retrieval | 0.65 |
| Average | 0.77 |
It works with the dependencies included in the latest version of MTEB.
TBD (technical report expected September 2025)