Downloads · 30 days
14
3% of all-time downloads
ADS509/BERTweet-large-self-labeling
BERTweet-large-self-labeling is a text classification model from ADS509. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
3% of all-time downloads
All-time downloads
406
Public
Parameters
355M
4.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.4 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of vinai/bertweet-large a dataset consisting of social media comments from 5 separate sources It achieves the following results on the evaluation set:
We retrained the classification layer of Bert Base for a multi-label classification task on our self-labeled data. The model description of the base model can be found at the link above and the description of the dataset can be found here. The fine-tuning parameters are listed below. The initial model used in this experiment was bert-base-uncased. After decent results, we decided to use this model as it was pre-trained on a copious amount of Twitter data, which more closely aligned with our dataset. Turned out to be a good decision as this model was a 7.2% improvement over bert-base on the evaluation data.
Intended use for this model is to better understand the nature of different social media websites and the nature of the discourse on that site beyond the usual "positive", "negative", "neutral" sentiment of most models. The labels for the commentary data are as follows:
We think there is promise in this approach, and as this is the initial step towards a deeper understanding of social commentary, there are several limitations to outline
A full description of the dataset can be found here
The full code used for training is below. We found overfitting to occur after 2 epochs
tokenizer = AutoTokenizer.from_pretrained("bert-base_uncased")
# Function to tokenize data with
def tokenize_function(batch):
return tokenizer(
batch['text'],
truncation=True,
max_length=512 # Can't be greater than model max length
)
# Tokenize Data
train_data = dataset['train'].map(tokenize_function, batched=True)
test_data = dataset['test'].map(tokenize_function, batched=True)
valid_data = dataset['valid'].map(tokenize_function, batched=True)
# Convert lists to tensors
train_data.set_format("torch", columns=['input_ids', "attention_mask", "label"])
test_data.set_format("torch", columns=['input_ids', "attention_mask", "label"])
valid_data.set_format("torch", columns=['input_ids', "attention_mask", "label"])
model = AutoModelForSequenceClassification.from_pretrained(
MODEL_ID,
num_labels=5, # adjust this based on number of labels you're training on
device_map='cuda',
dtype='auto',
label2id=label2id,
id2label=id2label
)
# Metric function for evaluation in Trainer
def compute_metrics(eval_pred):
predictions, labels = eval_pred
predictions = np.argmax(predictions, axis=1)
return {
'accuracy': accuracy_score(labels, predictions),
'f1_macro': f1_score(labels, predictions, average='macro'),
'f1_weighted': f1_score(labels, predictions, average='weighted')
}
# Data collator to handle padding dynamically per batch
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)
training_args = TrainingArguments(
output_dir='./bert-comment',
num_train_epochs=2,
per_device_train_batch_size=32,
per_device_eval_batch_size=64,
learning_rate=2e-5,
weight_decay=0.01,
warmup_steps=300,
# Evaluation & saving
eval_strategy='epoch',
save_strategy='epoch',
load_best_model_at_end=True,
metric_for_best_model='f1_macro',
# Logging
logging_steps=100,
report_to='tensorboard',
# Other
seed=42,
fp16=torch.cuda.is_available(), # Mixed precision if GPU available
)
# Set up Trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_data,
eval_dataset=valid_data,
processing_class=tokenizer,
data_collator=data_collator,
compute_metrics=compute_metrics
)
# Train!
trainer.train()
# Evaluate
eval_results = trainer.evaluate()
print(eval_results)
The following hyperparameters were used during training:
As this is a multi-label classification problem and there is class imbalance, the main metric we evaluate this model by is f1_macro
| Training Loss | Epoch | Step | Validation Loss | Accuracy | F1 Macro | F1 Weighted |
|---|---|---|---|---|---|---|
| 0.5943 | 1.0 | 1540 | 0.5735 | 0.7708 | 0.7592 | 0.7708 |
| 0.3951 | 2.0 | 3080 | 0.5607 | 0.7885 | 0.7817 | 0.7885 |