Downloads · 30 days
6
10% of all-time downloads
isomer007/takemeter-distilbert
takemeter-distilbert is a text classification model from isomer007. Use it when you need a label for a piece of text. It is set up for transformers.
isomer007/takemeter-distilbert is a 3-class text classifier fine-tuned from distilbert-base-uncased on 153 hand-labeled r/CryptoCurrency posts. It sorts crypto-community posts by the quality of the take — not by wheth…
Downloads · 30 days
6
10% of all-time downloads
All-time downloads
63
Public
Parameters
67M
268 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
isomer007/takemeter-distilbertisomer007/takemeter-distilbert is a 3-class text classifier fine-tuned from distilbert-base-uncased on 153 hand-labeled r/CryptoCurrency posts. It sorts crypto-community posts by the quality of the take — not by whether a post is bullish or bearish, but by whether its claim is backed by load-bearing, checkable evidence.
The fine-tuned model is the center of the AlphaSift project, which also includes a zero-shot llama-3.3-70b-versatile baseline and a Gradio demo. The complete writeup — taxonomy, label definitions, error analysis, and a side-by-side comparison with the baseline — lives in the project README.
The headline finding from the project evaluation is that this fine-tuned model loses to a zero-shot llama-3.3-70b-versatile baseline on the locked 23-post test set (0.522 vs 0.783 accuracy, a 0.261 gap). The cause is a data-quantity failure on 107 training examples, not a pipeline bug — see the Evaluation section below. The model is published because (a) the failure mode is itself a useful finding, and (b) classifier.py and app.py in the repo load it cleanly so the result is reproducible.
DistilBertForSequenceClassification (3-class single-label classification)distilbert-base-uncased)distilbert-base-uncasedpython app.py (Gradio) — see the project READMEThe model's intended use is classifying a single short Reddit-style post (≤256 tokens) as one of three labels:
| Label | Definition |
|---|---|
signal | Claim backed by load-bearing, verifiable evidence — on-chain data, a named mechanism, a concrete historical comparison. Direction (bullish or bearish) does not matter. |
hype | Bullish claim with no real evidence — bare price targets, FOMO, "to the moon." |
panic | Bearish claim with no real evidence — doom, "scam," nothing behind it. (The community's native term for this is "FUD.") |
It is a research / educational artifact that demonstrates the failure mode of fine-tuning a small encoder on a tiny hand-labeled dataset, and the value of comparing against a strong zero-shot baseline.
signal probability as a "quality score" attached to each post and aggregate it across a feed.filter-like content to this model will produce a forced 3-way prediction that should be ignored.max_length=256 at inference time; longer posts will have their tail silently cut off.llama-3.3-70b-versatile zero-shot on the same 23 posts. The fine-tuned model is not the recommended classifier in the AlphaSift project; the baseline is.hype) is well below the threshold at which a 6-layer transformer can pin down a reproducible decision boundary. The model has learned weak proxies — see the Reflection section in the project README — not the intended rule.hype), so a single flipped prediction moves that class's F1 by ~0.10–0.15. Per-class numbers should be read as directional, not precise.panic because of negative-coded vocabulary.signal. The "buying at 60k, betting on 200k" post — dense with price targets but no mechanism — is misread as signal instead of hype. This is the one error pattern that recurred identically across both training runs.signal. The dataset has more bullish-evidence posts than bearish-evidence posts, both because bullish-evidence content is more common on r/CryptoCurrency and because rigorous bearish analysis is harder to find. This contributes to the signal → panic error pattern.signal posts (e.g. "It is actually more profitable now to rent out GPU time to AI than it is to mine w/ ASICs") are underrepresented and frequently misread.Users (both direct and downstream) should be made aware of the risks, biases, and limitations of the model. More specific recommendations:
llama-3.3-70b-versatile zero-shot baseline from the same project, which outperforms it on the project's own test set.filter has been removed upstream), and only on short Reddit-style English text.signal posts, number-dense hype posts, balanced-sentiment signal posts), and (c) bootstrap confidence intervals on every reported metric, not point estimates on n=23.Use the code below to get started with the model. The transformers pipeline is the lightest path; the project's own classifier.py wraps the same call with the AlphaSift label order and a 256-token truncation.
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="isomer007/takemeter-distilbert",
top_k=None, # return probabilities for all 3 classes
truncation=True,
max_length=256,
)
result = classifier("9 UMA whale wallets control 53.1% of voting power. "
"Same wallets fund $12.5M Polymarket side-bets on "
"markets they resolve. The oracle is mathematically "
"incentivized to lie.")
# Returns: [[{'label': 'signal', 'score': ~0.34},
# {'label': 'panic', 'score': ~0.33},
# {'label': 'hype', 'score': ~0.33}]]
# Note: the project evaluates this exact post as correctly labeled 'signal',
# but the model's confidence is at chance — see the Limitations section.
# Or, for a single top label and confidence:
top = classifier("SOL is so back. This is the floor, screenshot this. LFG 🚀🚀")
print(top[0][0]) # {'label': 'hype', 'score': ~0.34}
For the full project — including the Groq baseline, evaluation, and Gradio demo — clone the repo and use the included classifier.py and app.py:
git clone https://github.com/isomer04/alphasift
cd alphasift
pip install -r requirements.txt
python app.py # Gradio demo at http://localhost:7860
python evaluate.py # reproduces the locked 23-post test-set evaluation
filter rows and stratified-splitting 70/15/15), drawn from a hand-collected pool of ~250 r/CryptoCurrency posts.filter rate of 25.7% (53/206). The remaining 153 are split:
signal: 65 rows (42.5% of usable)panic: 54 rows (35.3% of usable)hype: 34 rows (22.2% of usable)data/alphasift_dataset.csv (with [REVIEWED] flags in the notes column showing which rows were hand-corrected from the AI prelabel).data/taxonomy.md in the project repo.prelabeled in the notes column).[REVIEWED] in notes, with the original prelabel preserved in the note text where it was overridden along with the reasoning for the change.distilbert-base-uncased (6 layers, 768 hidden dim, 12 attention heads, 30522 vocab, tied input/output embeddings).id2label = {0: "hype", 1: "panic", 2: "signal"}, taken from config.py so the head's label order matches the rest of the repo.max_length=256, and padded. filter rows are dropped before training, so the model only ever sees signal / hype / panic.Trainer with the following hyperparameters:| Hyperparameter | Value |
|---|---|
| Epochs | 3 |
| Learning rate | 2e-5 |
| Train batch size | 16 |
| Weight decay | 0.01 |
| Warmup steps | 50 |
| Max sequence length | 256 |
| Best-model selection | highest validation accuracy (load_best_model_at_end) |
| Random seed | 42 |
random_state=42) so the test set is reproducible locally via evaluate.py.All numbers below are on the same locked 23-post test set (the 15% test split from the stratified split above). The full per-class breakdown, confusion matrix, and per-post error analysis are in the project README's Evaluation report section.
signal: 10 postspanic: 8 postshype: 5 postsmax_length=256.python evaluate.py regenerates the exact 23 posts from the locked seed and re-runs both models.The evaluation is disaggregated by:
signal / panic / hype) — reported as per-class precision, recall, F1, and support, plus a confusion matrix.signal → panic for bearish-evidence posts (the model uses negative-sounding vocabulary as a proxy for panic).signal.The evaluation is not disaggregated by subreddit, language, or time period (all 23 posts are r/CryptoCurrency, English, contemporaneous with the training data).
classification_report output.signal → panic, 6 of 10 true signal posts) visible at a glance.| Class | Precision | Recall | F1 | Support |
|---|---|---|---|---|
hype | 0.67 | 0.40 | 0.50 | 5 |
panic | 0.43 | 0.75 | 0.55 | 8 |
signal | 0.67 | 0.40 | 0.50 | 10 |
| macro avg | 0.59 | 0.52 | 0.52 | 23 |
| weighted avg | 0.58 | 0.52 | 0.52 | 23 |
Overall accuracy: 0.522
Confusion matrix (rows = true, columns = predicted):
| True ↓ \ Pred → | hype | panic | signal | Total |
|---|---|---|---|---|
hype | 2 | 2 | 1 | 5 |
panic | 1 | 6 | 1 | 8 |
signal | 0 | 6 | 4 | 10 |
| Total predicted | 3 | 14 | 6 | 23 |
llama-3.3-70b-versatile via Groq, no fine-tuning)For context — the baseline that beat this model on the same test set:
| Class | Precision | Recall | F1 | Support |
|---|---|---|---|---|
hype | 0.67 | 0.80 | 0.73 | 5 |
panic | 0.78 | 0.88 | 0.82 | 8 |
signal | 0.88 | 0.70 | 0.78 | 10 |
| macro avg | 0.77 | 0.79 | 0.78 | 23 |
| weighted avg | 0.80 | 0.78 | 0.78 | 23 |
Overall accuracy: 0.783
The zero-shot baseline wins on every class: hype 0.73 vs 0.50 F1, panic 0.82 vs 0.55 F1, signal 0.78 vs 0.50 F1. It never saw a single training example and still outperforms this model, because it brings pretrained world knowledge of what each kind of take sounds like — knowledge that 107 examples can't replace.
| Model | Test accuracy |
|---|---|
Zero-shot baseline (llama-3.3-70b-versatile) | 0.783 |
| Fine-tuned DistilBERT (this model) | 0.522 |
| Difference | −0.261 (this model regresses vs. baseline) |
The dominant error pattern is signal → panic (6 of 10 true signal posts). The model is not learning the intended evidence-quality rule; it is learning two cheaper proxies — negative-sounding words ⇒ panic and numeric density ⇒ signal/non-panic. The full per-post error analysis is in the project README's "Wrong predictions, analyzed" section.
The model also exhibits run-to-run instability: re-running the same training notebook on the same data with the same seed produced a different model where hype collapsed to 0% recall instead of signal. The specific failure mode is not stable; the underlying problem (data quantity) is.
The model's learned behavior was examined qualitatively via per-post error analysis on the locked test set, not via internal interpretability work (attention probing, probing classifiers, etc.). Three findings:
panic, even when the post is rigorously evidence-backed (the FTX/Anthropic stake breakdown is the clearest example). This is consistent with the model having recovered a sentiment axis it was never asked to learn.signal, even when the numbers are bare price targets with no mechanism (the "buying at 60k, betting on 200k" post). This proxy is the one error pattern that recurred identically across both training runs.No formal probing, attention visualization, or embedding-space analysis has been performed; these would be reasonable follow-ups but are out of scope for the v1.
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
llama-3.3-70b-versatile calls via Groq), which is the recommended classifier and is not done with this model at all.DistilBertForSequenceClassification — a 6-layer DistilBERT encoder with a 3-class linear classification head on top of the [CLS] pooled output. seq_classif_dropout=0.2 on the pooled output before the head.problem_type: "single_label_classification").{hype, panic, signal}. The head's id2label is {0: "hype", 1: "panic", 2: "signal"} and label2id is {hype: 0, panic: 1, signal: 2}. Note: the project repo's config.py LABEL_MAP matches this order exactly; classifier.py verifies the loaded model's id2label matches config.ID_TO_LABEL at load time and raises if not.transformers_version: 5.10.2 in config.json)gradio, scikit-learn, pandas, numpy — see requirements.txt in the project repoIf you use this model, please cite the AlphaSift project:
APA:
isomer007. (2026). AlphaSift: A fine-tuned DistilBERT text classifier for
crypto-community take-quality. Hugging Face Model ID: isomer007/takemeter-distilbert.
https://huggingface.co/isomer007/takemeter-distilbert
BibTeX:
@misc{alphasift2026takemeter,
author = {isomer007},
title = {AlphaSift: a fine-tuned DistilBERT text classifier for
crypto-community take-quality},
year = {2026},
howpublished = {\url{https://huggingface.co/isomer007/takemeter-distilbert}},
note = {Model card. Fine-tuned from \texttt{distilbert-base-uncased}
on 153 hand-labeled r/CryptoCurrency posts. Evaluated against
a zero-shot \texttt{llama-3.3-70b-versatile} baseline; the
baseline outperforms the fine-tuned model on the project's
locked 23-post test set (0.783 vs 0.522 accuracy). Project
repository: \url{https://github.com/isomer04/alphasift}.}
}
Base model citation (DistilBERT):
@inproceedings{sanh2019distilbert,
title = {DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter},
author = {Sanh, Victor and Debut, Lysandre and Chaumond, Julien and Wolf, Thomas},
booktitle = {EMC2 NLP Workshop at NeurIPS 2019},
year = {2019}
}
signal — a post whose claim is backed by load-bearing, checkable evidence (on-chain data, a named mechanism, a concrete historical comparison). Direction (bullish or bearish) does not matter.hype — a bullish post with no real evidence. The community's native term for the equivalent bearish case is "FUD"; this project calls that panic for symmetry.panic — a bearish post with no real evidence.filter — a post with no stance at all (news links, neutral questions, memes). Excluded from the 3-way taxonomy. Not a class of this model; the model is trained only on signal / hype / panic.panic in this taxonomy.hype.llama-3.3-70b-versatile via the Groq API.random_state=42, reproducible locally via python evaluate.py. The "locked" framing is to emphasize that all reported numbers in the project are on the same posts.data/taxonomy.mddocs/collecting.mddocs/planning.mdresults/confusion_matrix.pngisomer007
isomer007 on Hugging Face — open an issue on the AlphaSift project repo for the fastest response.