Downloads · 30 days
58
2% of all-time downloads
SandboxBhh/sentiment-thai-text-model
sentiment-thai-text-model is a text classification model from SandboxBhh. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
58
2% of all-time downloads
All-time downloads
3.3K
Public
Parameters
105M
3.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors421 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of poom-sci/WangchanBERTa-finetuned-sentiment on an pythainlp/wisesight_sentiment.
This model is a fine-tuned version of poom-sci/WangchanBERTa-finetuned-sentiment, specifically tailored for sentiment analysis on Thai-language texts. The fine-tuning was performed to improve performance on a custom Thai dataset for sentiment classification. The model is based on WangchanBERTa, a powerful transformer-based language model developed for Thai by the National Electronics and Computer Technology Center (NECTEC) in Thailand.
This model is designed to perform sentiment analysis, categorizing input text into three classes: positive, neutral, and negative. It can be used in a variety of natural language processing (NLP) applications such as:
Social media sentiment analysis Product or service reviews sentiment classification Customer feedback processing
Limitations: Language: The model is specialized for Thai text and may not perform well with other languages. Generalization: The model's performance depends on the quality and diversity of the dataset used for fine-tuning. It may not generalize well to domains that differ significantly from the training data. Ambiguity: Handling of highly ambiguous or sarcastic sentences may still be challenging.
The model was fine-tuned on a sentiment classification dataset composed of Thai-language text. The dataset includes sentences and texts from multiple domains, such as social media, product reviews, and general user feedback, labeled into three categories:
Positive: Indicates that the text expresses positive sentiment. Neutral: Indicates that the text is neutral or objective in sentiment. Negative: Indicates that the text expresses negative sentiment. More details on the dataset used can be provided upon request.
The model was trained using the following hyperparameters:
Learning rate: 2e-05 Batch size: 32 for both training and evaluation Seed: 42 (for reproducibility) Optimizer: Adam (with betas=(0.9, 0.999) and epsilon=1e-08) Scheduler: Linear learning rate scheduler Number of epochs: 5 The training used a combination of cross-entropy loss for multi-class classification and early stopping based on evaluation metrics.
The following hyperparameters were used during training: