Downloads · 30 days
5
26% of all-time downloads
a1a1a1aaa/structure-completeness-distilbert-zh
structure-completeness-distilbert-zh is a text classification model from a1a1a1aaa. Use it when you need a label for a piece of text. The card lists the license as other.
License: PolyForm Noncommercial 1.0.0 — personal learning / research / education / non-profit use only, no commercial use. This is a "source-available" model, not an OSI-approved open-source license (the restriction o…
Downloads · 30 days
5
26% of all-time downloads
All-time downloads
19
Public
Parameters
135M
541 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors541 MB · 99%
From the Hugging Face model README
License: PolyForm Noncommercial 1.0.0 — personal learning / research / education / non-profit use only, no commercial use. This is a "source-available" model, not an OSI-approved open-source license (the restriction on commercial use is the reason). See the license link for the full text before using this model in any product or paid service.
5-class classifier that scores the structural completeness of a Chinese interview answer (does it follow a clear situation/action/result-style structure vs. being disorganized rambling), as one dimension of an AI mock interview coaching app's answer-scoring pipeline. It is not a general Chinese text classifier and was not trained for any other task.
Output is one of 5 ordinal bands, each corresponding to a 0-10 structural completeness score range:
| label id | band | meaning |
|---|---|---|
| 0 | 0-2 | little to no structure |
| 1 | 3-4 | minimal structure |
| 2 | 5-6 | partial structure |
| 3 | 7-8 | mostly complete structure |
| 4 | 9-10 | fully complete, clearly organized structure |
structure_completeness score
(0-10) from a single human reviewer (no second independent rater —
see "Known limitations" below).Held-out test set (n=22, one single predict() call, never touched during
training):
| metric | value |
|---|---|
| exact_accuracy | 0.773 |
| macro_f1 | 0.769 |
| qwk (quadratic-weighted kappa) | 0.936 |
| within1_accuracy (+/-1 band) | 1.000 |
For reference, the 5-fold cross-validation mean on the 128-sample CV pool (different data split, not directly comparable 1:1 to the 22-sample test set above) was macro_f1=0.885, qwk=0.950. The test-set macro_f1/exact_accuracy are noticeably lower than the CV mean, which is expected given n=22 is a small, high-variance sample; qwk and within1_accuracy stay close to the CV mean, and every miss on the test set was an adjacent-band miss (no error was off by 2+ bands).
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "a1a1a1aaa/structure-completeness-distilbert-zh"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
answer = "首先我发现了这个问题,然后评估了几个方案,最后选择了风险最低的一个并推动落地。"
inputs = tokenizer(answer, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
band_id = int(torch.argmax(logits, dim=-1))
band_labels = ['0-2', '3-4', '5-6', '7-8', '9-10']
print(f"predicted band: {band_labels[band_id]}")
Or with pipeline:
from transformers import pipeline
clf = pipeline("text-classification", model="a1a1a1aaa/structure-completeness-distilbert-zh")
print(clf("首先我发现了这个问题,然后评估了几个方案,最后选择了风险最低的一个并推动落地。"))
Trained as part of a 10-week solo AI mock interview coaching app project.
Full training/eval code, data prep, cross-validation splits, and the
decision log documenting model selection reasoning are in the project repo:
https://github.com/Daifanqi/ai-interview-coach (ml/train.py,
ml/common.py, docs/week7_finetuning_results.md, docs/decision_log.md
decisions #33-#34).