Downloads · 30 days
21
0% of all-time downloads
AfterRain007/cryptobertRefined
cryptobertRefined is a text classification model from AfterRain007. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
CryptoBERTRefined is a fine tuned model from CryptoBERT by Elkulako model.
Downloads · 30 days
21
0% of all-time downloads
All-time downloads
10.3K
Public
Parameters
125M
499 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
CryptoBERTRefined is a fine tuned model from CryptoBERT by Elkulako model.
Input:
!pip -q install transformers
from transformers import TextClassificationPipeline, AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("AfterRain007/cryptobertRefined", use_fast=True)
model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels = 3)
pipe = TextClassificationPipeline(model=model, tokenizer=tokenizer, max_length=128, truncation=True, padding = 'max_length')
post_3 = "Because Forex Markets have years of solidity and millions in budget, not to mention that they use their own datacenters. These lame cryptomarkets are all supported by some Amazon-cloud-style system. They delegate and delegate their security and in the end, get buttfucked..."
post_2 = "Russian crypto market worth $500B despite bad regulation, says exec https://t.co/MZFoZIr2cN #CryptoCurrencies #Bitcoin #Technical Analysis"
post_1 = "I really wouldn't be asking strangers such an important question. I'm sure you'd get well meaning answers but you probably need professional advice."
df_posts = [post_1, post_2, post_3]
preds = pipe(df_posts)
print(preds)
Output:
[{'label': 'Neutral', 'score': 0.8427615165710449}, {'label': 'Bullish', 'score': 0.5444369912147522}, {'label': 'Bearish', 'score': 0.8388379812240601}]
Total of 3.803 text have been labelled manually to fine tune the model, with consideration of non-duplicate and a minimum of 4 words after cleaning. The following website were used for our training dataset:
Data augmentation was also performed to enrich the dataset, Back-Translation was used with Google Translate API on 10 language ('it', 'fr', "sv", "da", 'pt', 'id', 'pl', 'hr', "bg", "fi").
See Github for the source code to finetune cryptoBERT model into cryptoBERTRefined.
Credit where credit is due, thank you for all!