Downloads · 30 days
0
akaruineko/git-commits-sorter
git-commits-sorter is a text classification model from akaruineko. Use it when you need a label for a piece of text. The card lists the license as mit.
Classify raw git commit messages into categories (bugfix, feature, chore, docs, refactor, etc.) using a hybrid ML pipeline: TF-IDF + metadata → SVM / RF / CatBoost.
Downloads · 30 days
0
Access
Public
Updated Jul 20, 2026
Repo size
3.2 MB
Likes
0
Public
Click a slice to open those files.
.joblib3.2 MB · 100%
From the Hugging Face model README
Classify raw git commit messages into categories (bugfix, feature, chore, docs, refactor, etc.) using a hybrid ML pipeline: TF-IDF + metadata → SVM / RF / CatBoost.
--gpu flag enables cuML / CatBoost GPU / XGBoost GPU automaticallyparser.py → GitHub API → commits.jsonl (resilient scraping)
build_dataset.py → commits.jsonl → dataset.jsonl (auto-label + feature extraction)
train.py → dataset.jsonl → model.joblib (train & save)
predict.py → model.joblib → label (inference CLI)
commitsorter.py → CommitClassifier class (library API)
# single prediction
python predict.py "fix broad phase crash" 3 45 12
# positional args: message [files_count] [additions] [deletions]
python predict.py "bump lodash to 2.1"
# batch via stdin
echo '{"text":"fix crash","files_count":2}
{"text":"add login page","files_count":5,"additions":200}' | python predict.py --stdin
from commitsorter import CommitClassifier
clf = CommitClassifier("model.joblib")
label, scores = clf.sort("fix renderer crash", files_count=2, additions=10)
# → ("bugfix", [("bugfix", 0.77), ("chore", 0.11), ...])
results = clf.sort_batch([
{"text": "update dependencies", "files_count": 1},
{"text": "add user auth", "files_count": 5, "additions": 200},
])
# single model (default: SVM)
python train.py svm
python train.py lr
python train.py rf
python train.py cb
python train.py xgb
# tournament — compare all models
python train.py --all
# GPU acceleration (requires cuML / CUDA)
python train.py svm --gpu
python train.py --all --gpu
Each training sample:
| Field | Type | Description |
|---|---|---|
text | string | Commit message (prefix stripped) |
files_count | int | Number of files changed |
additions | int | Lines added |
deletions | int | Lines deleted |
changed_tests | int | Test files touched |
changed_docs | int | Doc files touched |
changed_source | int | Source files touched |
has_tests | bool | Any file in a test directory |
has_docs | bool | Any file in a docs directory |
extensions | [string] | Unique file extensions |
directories | [string] | Unique top-level directories |
label | string | Target class |
commit message ──→ TF-IDF(10k unigrams+bigrams)
file metadata ──→ StandardScaler
extensions ──→ OneHotEncoder
directories ──→ OneHotEncoder
↓
ColumnTransformer
↓
LinearSVC / RF / CatBoost / XGBoost
Best result: LinearSVC at 79% accuracy on held-out test set.
| Label | Heuristics |
|---|---|
bugfix | fix:, :bug:, "crash", "prevent" |
feature | feat:, :sparkles:, "add", "implement" |
chore | chore:, "bump", "update", "lint" |
docs | docs:, file-only docs/, .md |
refactor | refactor:, "cleanup", "remove", "rename" |
test | test:, file-only spec/ or test/ |
build | build:, "upgrade", dependency changes |
ci | ci:, CI config changes |
perf | perf:, "optimize" |
style | style:, "format", "prettier" |
revert | revert: |
# needs GITHUB_TOKEN in .env
python parser.py
Collects commits from configured repos into commits.jsonl. Resumes on interrupt via progress.json.