Downloads · 30 days
118
24% of all-time downloads
imocha-ai-org/ssf-skill-extractor
ssf-skill-extractor is a sentence similarity model from imocha-ai-org. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
A sentence-transformers model fine-tuned from all-MiniLM-L6-v2 for matching job description sentences to standardized skills from Singapore's SkillsFuture Framework (SSF).
Downloads · 30 days
118
24% of all-time downloads
All-time downloads
489
Public
Parameters
22.7M
90.9 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors90.9 MB · 99%
From the Hugging Face model README
A sentence-transformers model fine-tuned from all-MiniLM-L6-v2 for matching job description sentences to standardized skills from Singapore's SkillsFuture Framework (SSF).
The model maps sentences and skill names into a 384-dimensional dense vector space where job description text lands close to its corresponding skill, enabling accurate semantic skill extraction, tagging, and retrieval.
all-MiniLM-L6-v2 — same API, better skill matching| Property | Value |
|---|---|
| Model Type | Sentence Transformer (Bi-Encoder) |
| Base Model | sentence-transformers/all-MiniLM-L6-v2 |
| Architecture | BERT (6 layers, 12 heads, 384 hidden) |
| Parameters | ~22M |
| Max Sequence Length | 256 tokens |
| Output Dimensionality | 384 |
| Similarity Function | Cosine Similarity |
| Pooling | Mean Pooling + L2 Normalization |
| Language | English |
| License | Apache 2.0 |
SentenceTransformer(
(0): Transformer({'max_seq_length': 256, 'do_lower_case': False, 'architecture': 'BertModel'})
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_mean_tokens': True})
(2): Normalize()
)
| Property | Value |
|---|---|
| Name | SSF Skill Extraction Pairs |
| Domain | Workforce Skills / HR / Job Descriptions |
| Source Skills | 2,196 unique skills from Singapore SkillsFuture Framework |
| Synthetic Sentences | 5 JD-style sentences per skill, generated via Qwen3-1.7B (Ollama) |
| Total Training Pairs | 21,958 (positive + hard negative per sentence) |
| Format | (sentence, skill_name, label) — label 1.0 for correct skill, 0.0 for random incorrect skill |
| Validation Split | 10% held-out (2,195 pairs) |
Sample training pairs:
| Sentence | Skill | Label |
|---|---|---|
| Analyzes tax liabilities, identifies applicable rates, and applies corrections to ensure proper calculation and reporting. | Tax Computation | 1.0 |
| Monitor plant health by assessing symptoms and identifying disease risks. | Plant Health Management and Disease Control | 1.0 |
| Analyzes cross-cultural communication challenges in medical and legal contexts, optimizing translation strategies for diverse stakeholders. | Audience Segmentation | 0.0 |
Loss Function: CosineSimilarityLoss with MSE
The model learns to maximize cosine similarity between a JD sentence and its correct skill, while minimizing similarity to randomly-sampled incorrect skills. This contrastive setup produces well-separated embeddings.
| Parameter | Value |
|---|---|
| Epochs | 5 |
| Batch Size | 64 |
| Learning Rate | 5e-05 |
| Optimizer | AdamW (fused) |
| Warmup Steps | 10% of total steps |
| Scheduler | Linear decay |
| Seed | 42 |
| Precision | FP32 |
| Deterministic | Yes (CUBLAS_WORKSPACE_CONFIG=:4096:8) |
| Epoch | Step | Training Loss |
|---|---|---|
| 1.45 | 500 | 0.0822 |
| 2.91 | 1,000 | 0.0567 |
| 4.36 | 1,500 | 0.0493 |
Embeddings encoded with normalize_embeddings=True. Cosine similarity computed as dot product of normalized vectors.
| Model | AUC | Acc @ 0.5 | Best Accuracy | Pos Mean Sim | Neg Mean Sim |
|---|---|---|---|---|---|
| all-MiniLM-L6-v2 (baseline) | 0.978 | 0.810 | 0.928 | 0.530 | 0.133 |
| SSF-MiniLM v1 (1 epoch) | 0.989 | 0.949 | 0.952 | 0.799 | 0.131 |
| SSF-MiniLM v2 (5 epochs) | 0.995 | 0.968 | 0.971 | 0.845 | 0.088 |
pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer("imocha-ai-org/ssf-miniLM-finetuned-v2")
# Encode job description sentences and skills
sentences = [
"Design and implement scalable data pipelines for real-time analytics.",
"Manage patient records and ensure compliance with healthcare regulations.",
]
skills = [
"Data Engineering",
"Healthcare Records Management",
"Polymer Processing",
]
sentence_embeddings = model.encode(sentences, normalize_embeddings=True)
skill_embeddings = model.encode(skills, normalize_embeddings=True)
# Compute similarity (dot product of normalized vectors = cosine similarity)
import numpy as np
similarities = np.dot(sentence_embeddings, skill_embeddings.T)
print(similarities)
# sentence 0 -> "Data Engineering" = high score
# sentence 1 -> "Healthcare Records Management" = high score
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("imocha-ai-org/ssf-miniLM-finetuned-v2")
# Your skill taxonomy (or load from SSF dataset)
skills = ["Data Engineering", "Machine Learning", "Project Management", "Cloud Computing"]
skill_embeddings = model.encode(skills, normalize_embeddings=True)
# Extract skills from a JD sentence
jd_sentence = "Build and deploy ML models on AWS with CI/CD pipelines."
jd_embedding = model.encode([jd_sentence], normalize_embeddings=True)
scores = np.dot(jd_embedding, skill_embeddings.T)[0]
threshold = 0.5
for skill, score in sorted(zip(skills, scores), key=lambda x: -x[1]):
if score >= threshold:
print(f" {skill}: {score:.3f}")
from transformers import AutoTokenizer, AutoModel
import torch
tokenizer = AutoTokenizer.from_pretrained("imocha-ai-org/ssf-miniLM-finetuned-v2")
model = AutoModel.from_pretrained("imocha-ai-org/ssf-miniLM-finetuned-v2")
def encode(texts):
inputs = tokenizer(texts, padding=True, truncation=True, max_length=256, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
# Mean pooling
attention_mask = inputs["attention_mask"].unsqueeze(-1)
embeddings = (outputs.last_hidden_state * attention_mask).sum(1) / attention_mask.sum(1)
# L2 normalize
return torch.nn.functional.normalize(embeddings, p=2, dim=1)
query = encode(["Build scalable APIs with microservice architecture"])
skills = encode(["API Development", "Microservice Architecture", "Gardening"])
similarities = torch.mm(query, skills.T)
print(similarities)
| Property | Detail |
|---|---|
| Model Size | ~87 MB (safetensors) |
| Inference Speed | ~5,000 sentences/sec on GPU, ~500/sec on CPU (batch 64) |
| Memory | ~350 MB RAM loaded |
| ONNX Compatible | Yes (via sentence-transformers export) |
| Quantization | Compatible with INT8/FP16 for faster inference |
| Recommended Hardware | Works on CPU; GPU recommended for batch processing |
| Serving | Compatible with Triton, TorchServe, FastAPI, or any ONNX runtime |
The training dataset is available at imocha-ai-org/ssf-skill-extraction-pairs and contains:
pairs.jsonl — 21,958 training pairs (sentence, skill, label)generated_sentences.json — 5 synthetic JD sentences per skill (2,196 skills)meta.json — dataset metadata@misc{imocha2026ssf-miniLM,
title = {SSF-MiniLM Finetuned v2: Skill Extraction Embedding Model},
author = {imocha AI},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/imocha-ai-org/ssf-miniLM-finetuned-v2}
}
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}