Downloads · 30 days
283
50% of all-time downloads
fredoline005/ajan-embed-q
ajan-embed-q is a sentence similarity model from fredoline005. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
A Turkish-optimized sentence embedding model for retrieval / RAG, distilled from BAAI/bge-m3 into a multilingual-e5-base student over 500k Turkish web sentences. 768→1024-dim projected to match the teacher's space.
Downloads · 30 days
283
50% of all-time downloads
All-time downloads
569
Public
Parameters
278M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
A Turkish-optimized sentence embedding model for retrieval / RAG, distilled from
BAAI/bge-m3 into a multilingual-e5-base student over 500k Turkish web sentences.
768→1024-dim projected to match the teacher's space.
🔗 Code + recipe: github.com/AJANLAR-AI/ajanlar
Part of Ajanlar — open Turkish AI agent infra. The retrieval engine under the agents, where multilingual models underperform on agglutinative Turkish.
Main score per task (NDCG@10 retrieval; Spearman STS). multilingual-e5-base is the
undistilled ablation (same base as this model).
| Model | Params | TurHistQuad (retrieval) | STS22.v2 | STS17 | avg |
|---|---|---|---|---|---|
| ajan-embed-q (this) | 278M | 0.465 | 0.651 | 0.724 | 0.613 |
| multilingual-e5-small | 118M | 0.433 | 0.643 | 0.767 | 0.614 |
| multilingual-e5-base (ablation) | 278M | 0.444 | 0.651 | 0.777 | 0.624 |
| multilingual-e5-large | 560M | 0.469 | 0.675 | 0.810 | 0.652 |
| BAAI/bge-m3 (teacher) | 568M | 0.478 | 0.680 | 0.814 | 0.657 |
Honest reading: this is a retrieval-specialised model. On Turkish retrieval (TurHistQuad) it beats e5-small/e5-base and matches e5-large at half the size — the RAG/agent use case it's built for. On general STS it trails the e5 family, so on the 4-task average it lands ~on par with (slightly below) undistilled e5-base. Use it for retrieval/RAG in Turkish, not as a general-purpose STS model.
from sentence_transformers import SentenceTransformer
m = SentenceTransformer("fredoline005/ajan-embed-q")
emb = m.encode(["query: kargom ne zaman gelir?",
"passage: Siparişler 1–3 iş günü içinde kargoya verilir."],
normalize_embeddings=True)
Use query: / passage: prefixes for retrieval (inherited from the e5 family).
BAAI/bge-m3 (MIT). Student: intfloat/multilingual-e5-base (MIT).allenai/c4, tr).Apache-2.0 (weights/recipe). Base + teacher are MIT.
The Ajanlar project is for sale as a whole: the ajanlar.ai domain, this model and
its training recipe, and the brand.
It was built because multilingual embedding models underperform on agglutinative
Turkish. A 278M-parameter distilled student closes most of that gap on retrieval and
matches multilingual-e5-large at half the size on TurHistQuad. The model is installed
roughly 240 times a month and that rate has about doubled since June with no promotion
of any kind.
This changes nothing for anyone using the model. It stays Apache-2.0, it stays up, and it stays free, whoever ends up owning it. If you are building on it, keep building.
Enquiries: [email protected]
Built by Ranked Technologies Limited (RC 9522220) as the retrieval engine under Ajanlar, Turkish-language AI support agents that answer from a business's own documents.
If you are using this model, we would genuinely like to know what for. If it is falling short on your data, tell us and we will look at it.
Contact: [email protected]
We also build Nigerian-language AI at 9jatesters and publish speech datasets and an open benchmark at 9jaBench.