Downloads · 30 days
30
31% of all-time downloads
hadangvu/pkd-albert-student
pkd-albert-student is a text classification model from hadangvu. Use it when you need a label for a piece of text. The card lists the license as mit.
[](https://huggingface.co/spaces/hadangvu/pkd-sentiment-api) [](https://huggingface.co/spaces/hadangvu/pkd-sentiment-api)
Downloads · 30 days
30
31% of all-time downloads
All-time downloads
97
Public
Parameters
11.7M
46.7 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors46.7 MB · 96%
From the Hugging Face model README
PKD-ALBERT is a lightweight financial sentiment classifier distilled from ProsusAI/finbert using a two-stage Patient Knowledge Distillation (PKD) pipeline.
It classifies financial text — headlines, earnings excerpts, news snippets — into positive, neutral, or negative sentiment, achieving 97.4% accuracy and 0.965 Macro-F1 on Financial PhraseBank while using ~10× fewer parameters than the teacher model.
<table> <thead> <tr> <th></th> <th>Teacher (FinBERT)</th> <th>Student (PKD-ALBERT)</th> </tr> </thead> <tbody> <tr> <td><strong>Parameters</strong></td> <td>109.5M</td> <td>11.7M</td> </tr> <tr> <td><strong>Model size</strong></td> <td>417.7 MB</td> <td>44.6 MB</td> </tr> <tr> <td><strong>Test Accuracy</strong></td> <td>97.6%</td> <td><strong>97.4%</strong></td> </tr> <tr> <td><strong>Macro F1</strong></td> <td>0.9696</td> <td><strong>0.9650</strong></td> </tr> <tr> <td><strong>Inference (ms/doc)</strong></td> <td>1.49 ms</td> <td>2.02 ms</td> </tr> <tr> <td><strong>Accuracy drop</strong></td> <td>—</td> <td><strong>−0.2%</strong></td> </tr> </tbody> </table>The student retains 99.8% of teacher accuracy at 89% less disk space.
A live Gradio demo is available on Hugging Face Spaces. Paste any financial headline or sentence and receive a sentiment label with a confidence score.
Example inputs from the held-out test set:
<table> <thead> <tr> <th>Input</th> <th>Prediction</th> <th>Confidence</th> </tr> </thead> <tbody> <tr> <td><em>"Charles Schwab price target raised to $121 from $119 at JPMorgan."</em></td> <td>✅ Positive</td> <td>93.1%</td> </tr> <tr> <td><em>"Costco assumed with a Peer Perform at Wolfe Research."</em></td> <td>➖ Neutral</td> <td>92.2%</td> </tr> <tr> <td><em>"IBM Explains How AI Models Are Making a Familiar Human Mistake."</em></td> <td>❌ Negative</td> <td>90.1%</td> </tr> </tbody> </table>The /predict endpoint accepts raw text and returns a label and confidence score.
import requests
API_URL = "https://hadangvu-pkd-sentiment-api.hf.space/predict"
response = requests.post(API_URL, json={
"text": "Charles Schwab price target raised to $121 from $119 at JPMorgan."
})
print(response.json())
# {
# "label": "positive",
# "confidence": 0.93,
# "latency_ms": 79.3
# }
Input: { "text": "..." } — any financial sentence or headline (max 128 tokens)
Output: { "label": str, "confidence": float, "latency_ms": float }
This model was trained using a two-stage Patient Knowledge Distillation strategy.
<h3>Stage 1 — Distillation on Pseudo-Labeled Financial News</h3>The student (ALBERT-base) was first trained on a large corpus of scraped financial news pseudo-labeled by the FinBERT teacher, using a combined loss of soft KL divergence targets and intermediate layer alignment.
<table> <thead> <tr> <th>Parameter</th> <th>Value</th> </tr> </thead> <tbody> <tr> <td>Dataset</td> <td>Scraped financial news (pseudo-labeled by FinBERT)</td> </tr> <tr> <td>Train / Val / Test split</td> <td>5,587 / 1,197 / 1,198</td> </tr> <tr> <td>Epochs</td> <td>3</td> </tr> <tr> <td>Batch size</td> <td>32</td> </tr> <tr> <td>Optimizer</td> <td>AdamW</td> </tr> <tr> <td>Learning rate</td> <td>2e-5</td> </tr> <tr> <td>KD temperature sweep</td> <td>[2, 5, 9]</td> </tr> <tr> <td>Alpha (KD loss weight)</td> <td>0.3</td> </tr> <tr> <td>PKD beta</td> <td>0.02</td> </tr> <tr> <td>PKD student layers</td> <td>[2, 4, 8, 12]</td> </tr> </tbody> </table> <h3>Stage 2 — Fine-tuning on Financial PhraseBank</h3>The distilled student was then fine-tuned on the high-quality Financial PhraseBank (100% annotator agreement) subset using standard cross-entropy loss to align the student with gold-label financial sentiment.
<table> <thead> <tr> <th>Parameter</th> <th>Value</th> </tr> </thead> <tbody> <tr> <td>Dataset</td> <td>Financial PhraseBank (100% agreement)</td> </tr> <tr> <td>Train / Val / Test split</td> <td>1,584 / 340 / 340</td> </tr> <tr> <td>Epochs</td> <td>1</td> </tr> <tr> <td>Loss</td> <td>Cross-entropy</td> </tr> </tbody> </table> <h3>Loss Function</h3>The Stage 1 total loss combines:
<ul> <li><strong>KL Divergence</strong> between teacher soft targets and student logits (soft label transfer)</li> <li><strong>Patient KD</strong> alignment between intermediate ALBERT and FinBERT hidden layers</li> <li><strong>Alpha</strong> controls the balance between hard label CE loss and soft KD loss</li> </ul>The table below compares all training strategies evaluated against the same Financial PhraseBank test set (340 samples):
<table> <thead> <tr> <th>Model</th> <th>Params</th> <th>Size</th> <th>Test Acc</th> <th>Macro F1</th> <th>KL (teacher→student)</th> </tr> </thead> <tbody> <tr> <td>Teacher FinBERT</td> <td>109.5M</td> <td>417.7 MB</td> <td>97.6%</td> <td>0.9696</td> <td>—</td> </tr> <tr> <td>Fresh → FP (baseline)</td> <td>11.7M</td> <td>44.6 MB</td> <td>77.1%</td> <td>0.6126</td> <td>0.359</td> </tr> <tr> <td>CE-scraped → FP</td> <td>11.7M</td> <td>44.6 MB</td> <td>95.9%</td> <td>0.9392</td> <td>0.230</td> </tr> <tr> <td>KD-scraped → FP</td> <td>11.7M</td> <td>44.6 MB</td> <td>96.5%</td> <td>0.9514</td> <td>0.156</td> </tr> <tr> <td><strong>PKD-scraped → FP (ours)</strong></td> <td><strong>11.7M</strong></td> <td><strong>44.6 MB</strong></td> <td><strong>97.4%</strong></td> <td><strong>0.9650</strong></td> <td>0.188</td> </tr> </tbody> </table>Key takeaway: Patient KD achieves the highest Macro F1 among all student variants and closes to within 0.5% of the teacher — demonstrating that intermediate layer alignment significantly improves distillation quality beyond standard KD.
If you use this model or the distillation pipeline in your work, please cite:
@misc{pkd-albert-finbert,
author = {Ha Dang Vu},
title = {PKD-ALBERT: Lightweight Financial Sentiment via Patient Knowledge Distillation},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/hadangvu/pkd-albert-student}
}
<em>Built as part of a financial NLP research project exploring efficient model compression for domain-specific sentiment analysis.</em>