Downloads · 30 days
93
29% of all-time downloads
Misbahuddin/job-title-normalizer-e5-base
job-title-normalizer-e5-base is a sentence similarity model from Misbahuddin. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as mit.
A sentence encoder fine-tuned to normalize messy, multilingual job titles to a canonical occupation taxonomy (ESCO + ONET), framed and evaluated as retrieval: embed a noisy or foreign-language title (the query) and re…
Downloads · 30 days
93
29% of all-time downloads
All-time downloads
317
Public
Parameters
278M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
A sentence encoder fine-tuned to normalize messy, multilingual job titles to a canonical occupation taxonomy (ESCO + O*NET), framed and evaluated as retrieval: embed a noisy or foreign-language title (the query) and retrieve the closest canonical occupation label (the passage) from a fixed corpus of 4,055 occupations.
Live demo: https://job-title-normalizer-525186107937.us-central1.run.app (Cloud Run; may cold-start ~1 min) · Lighter variant: job-title-normalizer-e5-small
Why retrieval, not classification? Occupation taxonomies have thousands of classes and the label set evolves constantly. A nearest-neighbour retriever over embeddings generalizes to unseen labels and lets you swap the corpus without retraining a softmax head.
Held-out test set of 15,248 real ESCO/O*NET titles against a 4,055-occupation corpus. The split is by occupation (zero overlap between train/val/test occupations, asserted at build time), so every test occupation is unseen during training.
| Slice | Method | Recall@1 | Recall@5 | Recall@10 | MRR |
|---|---|---|---|---|---|
| Overall | BM25 (lexical) | 0.095 | 0.163 | 0.183 | 0.124 |
| Overall | zero-shot e5-base | 0.237 | 0.378 | 0.437 | 0.298 |
| Overall | this model | 0.381 | 0.572 | 0.645 | 0.463 |
| FR→EN | this model | 0.548 | 0.774 | 0.847 | 0.643 |
| DE→EN | this model | 0.536 | 0.792 | 0.852 | 0.642 |
Cross-lingual context: BM25 scores MRR 0.034 on the FR/DE→EN slice (a French query shares almost no tokens with an English canonical label); this model reaches 0.643 — lexical search structurally cannot do this task, and fine-tuning adds ~35% over the zero-shot base.
intfloat/multilingual-e5-base (mean pooling, query:/passage: prefixes, 768-dim).MultipleNegativesRankingLoss, scale 20 ≈ temperature 0.05), wrapped in MatryoshkaLoss (dims 768/512/256/128/64) so truncated embeddings stay usable.NoDuplicatesDataLoader to reduce in-batch false negatives. Trained on an RTX 4060 Laptop (8 GB) in ~46 min.tn_preset.json) so downstream tooling recovers the correct behavior.The e5 family needs asymmetric prefixes: encode input titles as query: … and canonical
labels as passage: …. Forgetting them silently degrades accuracy.
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("Misbahuddin/job-title-normalizer-e5-base")
canonicals = [
"passage: software developer",
"passage: data scientist",
"passage: nurse responsible for general care",
]
query = "query: Ingénieur logiciel" # French → English canonical
q = model.encode(query, normalize_embeddings=True)
c = model.encode(canonicals, normalize_embeddings=True)
scores = cos_sim(q, c)[0]
print(canonicals[int(scores.argmax())], float(scores.max()))
# passage: software developer ~0.79
Vectors are L2-normalized, so cosine == dot product — index canonical embeddings with FAISS
IndexFlatIP for production search. Matryoshka: you may truncate embeddings to 256/128/64
dims (then re-normalize) for cheaper indexes without re-encoding.
Calibrate an abstention threshold. Out-of-taxonomy queries still return a nearest neighbour — but at tellingly low scores (e.g. "RevOps", which has no ESCO/O*NET occupation, scores ~0.31 vs ~0.8 for true matches). Reject below a threshold tuned on your data.
ONET® is a trademark of USDOL/ETA. This model was produced using ONET data but is not endorsed by USDOL/ETA. ESCO is a service of the European Commission; this model is not endorsed by the Commission.