Downloads · 30 days
0
Sanjay1603/classification-xgb
classification-xgb is a text classification model from Sanjay1603. Use it when you need a label for a piece of text. It is set up for sentence-transformers. The card lists the license as mit.
This repository contains an XGBoost classification model designed to categorize question difficulty into three levels: Easy, Medium, and Hard.
Downloads · 30 days
0
Access
Public
Updated Jun 27, 2026
Repo size
5 MB
Likes
0
Public
Click a slice to open those files.
.joblib5 MB · 100%
From the Hugging Face model README
This repository contains an XGBoost classification model designed to categorize question difficulty into three levels: Easy, Medium, and Hard.
question_content, any multiple choice options (formatted as a pipe-separated string), and associated passage_text context.sentence-transformers/all-MiniLM-L6-v2 embedding engine.XGBClassifier) predicts the difficulty class from the embedding vector.difficulty_xgb_local_model.joblib: The trained XGBoost model binary.difficulty_label_encoder.joblib: Scikit-Learn LabelEncoder mapping classes to integers (['Easy', 'Hard', 'Medium']).You can load and use these assets directly using joblib and sentence-transformers:
import joblib
import json
import pandas as pd
import numpy as np
from sentence_transformers import SentenceTransformer
# 1. Load the model, label encoder and embedder
model = joblib.load("difficulty_xgb_local_model.joblib")
label_encoder = joblib.load("difficulty_label_encoder.joblib")
embedder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
# 2. Define single item payload
question = "The number of species having non-pyramidal shape is..."
options = ["CO3 2-", "SO3", "NO3-"]
# 3. Format inputs (match training pipeline schema)
options_text = " | ".join(options)
enriched_text = f"Question: {question} \n Options: {options_text}"
# 4. Generate embeddings and run inference
embedding = embedder.encode([enriched_text])
prediction_idx = model.predict(embedding)[0]
predicted_label = label_encoder.classes_[prediction_idx]
print("Predicted Difficulty:", predicted_label)