Downloads · 30 days
13
7% of all-time downloads
visolex/textcnn-hsd
textcnn-hsd is a text classification model from visolex. Use it when you need a label for a piece of text. The card lists the license as mit.
This model is a fine-tuned version of unknown on the ViHSD (Vietnamese Hate Speech Detection Dataset) for classifying Vietnamese text into three categories: CLEAN, OFFENSIVE, and HATE.
Downloads · 30 days
13
7% of all-time downloads
All-time downloads
180
Public
Repo size
43 MB
Likes
0
Public
Click a slice to open those files.
.bin21.5 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of unknown on the ViHSD (Vietnamese Hate Speech Detection Dataset) for classifying Vietnamese text into three categories: CLEAN, OFFENSIVE, and HATE.
322e-51002560.015005Model was trained on ViHSD (Vietnamese Hate Speech Detection Dataset) containing ~10,000 Vietnamese comments from social media.
The model was evaluated on test set with the following metrics:
0.83880.30410.76520.27960.3333from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Load model and tokenizer
model_name = "visolex/textcnn-hsd"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
model_name
)
# Classify text
text = "Văn bản tiếng Việt cần phân loại"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_label = torch.argmax(predictions, dim=-1).item()
# Label mapping
label_names = {
0: "CLEAN",
1: "OFFENSIVE",
2: "HATE"
}
print(f"Predicted label: {label_names[predicted_label]}")
print(f"Confidence scores: {predictions[0].tolist()}")
⚠️ Note for Vocab-based Models: This model (textcnn) uses custom vocabulary-based tokenization and does not include a Hugging Face tokenizer. You will need to implement custom tokenization or load a tokenizer from a compatible base model. The model expects word-level tokenized input.
This model is distributed under the MIT License.