Downloads · 30 days
10
2% of all-time downloads
Rendika/tweets-election-classification
tweets-election-classification is a text classification model from Rendika. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This repository contains a fine-tuned of indolem/indobertweet-base-uncased model for classifying tweets related to election topics. The model has been trained to categorize tweets into eight distinct classes, providin…
Downloads · 30 days
10
2% of all-time downloads
All-time downloads
448
Public
Parameters
111M
442 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors442 MB · 100%
From the Hugging Face model README
This repository contains a fine-tuned of indolem/indobertweet-base-uncased model for classifying tweets related to election topics. The model has been trained to categorize tweets into eight distinct classes, providing valuable insights into public opinion and discourse during election periods.
The model classifies tweets into the following categories:
| Encoded | Label |
|---|---|
| 0 | Demografi |
| 1 | Ekonomi |
| 2 | Geografi |
| 3 | Ideologi |
| 4 | Pertahanan dan Keamanan |
| 5 | Politik |
| 6 | Sosial Budaya |
| 7 | Sumber Daya Alam |
The following libraries were used for data processing, model training, and evaluation:
numpy, pandas, re, string, randommatplotlib.pyplot, seaborn, tqdm, plotly.graph_objs, plotly.express, plotly.figure_factoryPIL, wordcloudnltk, nlp_id, Sastrawi, tweet-preprocessortensorflow, keras, sklearn, transformers, torchThe dataset was split into training, validation, and test sets with the following proportions:
| Epoch | Train Loss | Train Accuracy | Validation Loss | Validation Accuracy |
|---|---|---|---|---|
| 1 | 0.9382 | 0.7167 | 0.7518 | 0.7671 |
| 2 | 0.5741 | 0.8229 | 0.7081 | 0.7931 |
| 3 | 0.3541 | 0.8958 | 0.7473 | 0.7953 |
The model is built using the TensorFlow and Keras libraries and employs the following architecture:
To use the model, ensure you have the required libraries installed. You can install them using pip:
pip install transformers
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("Rendika/tweets-election-classification")
model = AutoModelForSequenceClassification.from_pretrained("Rendika/tweets-election-classification")
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-classification", model="Rendika/tweets-election-classification")
The data was cleaned using the following steps:
Here's a sample code snippet to load and use the model:
import tensorflow as tf
from tensorflow.keras.models import load_model
import pandas as pd
# Load the trained model
model = load_model('path_to_your_model.h5')
# Preprocess new data
def preprocess_text(text):
# Include your text preprocessing steps here
pass
# Example usage
new_tweets = pd.Series(["Your new tweet text here"])
preprocessed_tweets = new_tweets.apply(preprocess_text)
# Tokenize and pad sequences as done during training
# ...
# Predict the class
predictions = model.predict(preprocessed_tweets)
predicted_classes = predictions.argmax(axis=-1)
The model was evaluated using the following metrics:
This fine-tuned model provides a robust tool for classifying election-related tweets into distinct categories. It can be used to analyze public sentiment and trends during election periods, aiding in better understanding and decision-making.
This project is licensed under the MIT License.
For any questions or feedback, please contact [me] at [[email protected]].