Downloads · 30 days
27
23% of all-time downloads
gr8monk3ys/resume-section-classifier
resume-section-classifier is a text classification model from gr8monk3ys. Use it when you need a label for a piece of text. The card lists the license as mit.
A fine-tuned DistilBERT model that classifies resume text sections into 8 categories. Designed for automated resume parsing pipelines where incoming text needs to be segmented and labeled by section type.
Downloads · 30 days
27
23% of all-time downloads
All-time downloads
117
Public
Parameters
67M
268 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
A fine-tuned DistilBERT model that classifies resume text sections into 8 categories. Designed for automated resume parsing pipelines where incoming text needs to be segmented and labeled by section type.
| Label | Description | Example |
|---|---|---|
education | Academic background, degrees, coursework, GPA | "B.S. in Computer Science, MIT, 2023, GPA: 3.8" |
experience | Work history, job titles, responsibilities | "Software Engineer at Google, 2020-Present. Built microservices..." |
skills | Technical and soft skills listings | "Python, Java, React, Docker, Kubernetes, AWS" |
projects | Personal or academic project descriptions | "Built a real-time analytics dashboard using React and D3.js" |
summary | Professional summary or objective statement | "Results-driven engineer with 5+ years of experience in..." |
certifications | Professional certifications and licenses | "AWS Certified Solutions Architect - Associate (2023)" |
contact | Name, email, phone, LinkedIn, location | "John Smith |
awards | Honors, achievements, recognition | "Dean's List, Phi Beta Kappa, Hackathon Winner (2022)" |
The model was trained on a synthetic dataset generated programmatically using data_generator.py. The generator uses:
Default configuration produces 1,920 examples (80 base examples x 3 variants x 8 categories), split 80/10/10 into train/validation/test sets with stratified sampling.
| Parameter | Value |
|---|---|
| Base model | distilbert-base-uncased |
| Max sequence length | 256 tokens |
| Epochs | 4 |
| Batch size | 16 |
| Learning rate | 2e-5 |
| LR scheduler | Cosine |
| Warmup ratio | 0.1 |
| Weight decay | 0.01 |
| Optimizer | AdamW |
| Early stopping | Patience 3 (on F1 macro) |
These figures are INDICATIVE only. They are measured on a held-out split of the model's own synthetic, self-generated data (192 examples, stratified) — not a real-world benchmark. Because the train and test sets are produced by the same template-based generator, they share vocabulary and structure, so these scores overstate how the model will perform on real resumes.
Measured on the published checkpoint (4 epochs, seed 42, 192-example stratified synthetic test split):
| Metric | Score (synthetic test set) |
|---|---|
| Accuracy | 1.000 |
| F1 (macro) | 1.000 |
| F1 (weighted) | 1.000 |
| Precision (weighted) | 1.000 |
| Recall (weighted) | 1.000 |
| Eval loss | 0.195 |
A perfect score here is evidence of the leak described above, not of quality: the test split is generated by the same templates as the training split.
To give the synthetic score some context, the checkpoint was run against eight hand-written snippets in the style of real resumes, one per section type, none drawn from the generator:
| Result | Count |
|---|---|
| Correct | 6 / 8 |
Both misses were low-confidence and collapsed toward the broader categories —
an experience bullet read as summary (0.39), and a projects bullet read
as skills (0.48). Treat ~0.75 on real text as the more realistic
expectation, and treat predictions under ~0.5 confidence as unreliable.
Note: exact metrics depend on the random seed and training run.
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="gr8monk3ys/resume-section-classifier",
)
result = classifier(
"Bachelor of Science in Computer Science, Stanford University, 2023. GPA: 3.9/4.0"
)
print(result)
# [{'label': 'education', 'score': 0.98}]
from inference import ResumeSectionClassifier
classifier = ResumeSectionClassifier("gr8monk3ys/resume-section-classifier")
resume_text = """
John Smith
[email protected] | (415) 555-1234 | San Francisco, CA
linkedin.com/in/johnsmith | github.com/johnsmith
SUMMARY
Experienced software engineer with 5+ years building scalable web applications
and distributed systems. Passionate about clean code and mentoring junior developers.
EXPERIENCE
Senior Software Engineer | Google | San Francisco, CA | Jan 2021 - Present
- Led migration of monolithic application to microservices architecture
- Mentored 4 junior engineers through code reviews and pair programming
- Reduced API latency by 40% through caching and query optimization
Software Engineer | Stripe | San Francisco, CA | Jun 2018 - Dec 2020
- Built payment processing APIs handling 10M+ transactions daily
- Implemented CI/CD pipeline reducing deployment time from hours to minutes
EDUCATION
B.S. in Computer Science, Stanford University (2018)
GPA: 3.8/4.0 | Dean's List
SKILLS
Languages: Python, Java, Go, TypeScript
Frameworks: React, Django, Spring Boot, gRPC
Tools: Docker, Kubernetes, AWS, PostgreSQL, Redis
"""
analysis = classifier.classify_resume(resume_text)
print(analysis.summary())
# Install dependencies
pip install -r requirements.txt
# Generate data and train
python train.py --epochs 4 --batch-size 16
# Train and push to Hub
python train.py --push-to-hub --hub-model-id gr8monk3ys/resume-section-classifier
# Inference
python inference.py --file resume.txt
python inference.py --text "Python, Java, React, Docker" --single
# Generate data with custom settings
python data_generator.py \
--examples-per-category 150 \
--augmented-copies 3 \
--output data/resume_sections.csv \
--print-stats
resume-section-classifier/
data_generator.py # Synthetic data generation with templates and augmentation
train.py # Full fine-tuning pipeline with HuggingFace Trainer
inference.py # Section splitting and classification API + CLI
requirements.txt # Python dependencies
README.md # This model card
This model is not intended for making hiring decisions. It is a text classification tool for structural parsing only.
Lorenzo Scaturchio (gr8monk3ys)