Downloads · 30 days
15
24% of all-time downloads
jamal-ibrahim/roberta-large-cv-detector
roberta-large-cv-detector is a text classification model from jamal-ibrahim. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
This model is a fine-tuned RoBERTa-Large transformer trained to detect whether a curriculum vitae (CV) was written by a human, generated by AI, or produced through a combination of both.
Downloads · 30 days
15
24% of all-time downloads
All-time downloads
62
Public
Parameters
355M
1.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.4 GB · 100%
From the Hugging Face model README
This model is a fine-tuned RoBERTa-Large transformer trained to detect whether a curriculum vitae (CV) was written by a human, generated by AI, or produced through a combination of both.
The model was trained using the dataset:
jamal-ibrahim/ai-cv-detection-dataset
The task is formulated as a three-class classification problem:
Label Meaning
human CV written entirely by a human mixed CV containing both human-written and AI-generated content ai_generated CV generated primarily by an AI system
The goal of this experiment was to explore whether modern transformer models can identify stylistic differences between human-written and AI-generated professional documents.
Base model:
roberta-large
RoBERTa-Large contains approximately 355 million parameters and is known to perform well on tasks involving linguistic style, authorship detection, and semantic classification.
Dataset used for training:
jamal-ibrahim/ai-cv-detection-dataset
Total dataset size: approximately 1500 CV documents.
The dataset was split using a stratified strategy to maintain class balance.
Split configuration:
Split Percentage
Train 70% Validation 15% Test 15%
Each class (human, mixed, ai_generated) is represented equally in each split.
Training parameters:
Parameter Value
Epochs 10 Batch size 16 Learning rate 5e-6 Optimizer AdamW Max sequence length 512 Evaluation strategy per epoch
Training was performed using the Hugging Face Trainer API.
Evaluation was performed on the held-out test set.
True / Predicted Human Mixed AI
Human 75 0 0 Mixed 0 74 1 AI Generated 0 0 75
Only one misclassification occurred in the test set.
Metric Value
Accuracy ~0.995 Weighted F1 ~0.996 Macro F1 ~0.995
The model demonstrates near-perfect classification performance on this dataset.
The extremely high classification accuracy suggests that the model successfully learned stylistic patterns that differentiate:
Several factors may contribute to this performance.
Large language models often produce documents with consistent structural patterns such as:
These patterns may create detectable stylistic signatures.
Human-written CVs often contain:
This variability may make human documents distinguishable from AI-generated ones.
Documents labeled as mixed appear to occupy an intermediate stylistic space between human and AI-generated text.
The single observed misclassification (mixed → AI-generated) suggests that documents heavily edited by AI may resemble fully generated content.
Although the results are strong, several limitations should be considered.
The dataset contains approximately 1500 samples, which is relatively small for training large transformer models.
High performance may partly reflect the limited diversity of the dataset.
If AI-generated CVs were produced using similar prompts or generation patterns, the model may learn those specific templates rather than generalizable AI-writing signals.
This model was trained specifically on CV-style documents and may not generalize well to other types of text such as essays, emails, or reports.
Future experiments could explore several directions:
Potential applications include:
The model should not be used as a definitive tool for determining authorship in real-world hiring decisions.
@model{roberta_large_cv_detector, title = {RoBERTa-Large CV AI Detection Model}, author = {Jamal Ibrahim}, year = {2026}, dataset = {jamal-ibrahim/ai-cv-detection-dataset}, publisher = {Hugging Face} }