Downloads · 30 days
18
7% of all-time downloads
justtherightsize/small-e-czech-multi-label-online-risks-cs
small-e-czech-multi-label-online-risks-cs is a feature extraction model from justtherightsize. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
This model is fine-tuned for multi-label text classification of Online Risks in Instant Messenger dialogs of Adolescents.
Downloads · 30 days
18
7% of all-time downloads
All-time downloads
247
Public
Repo size
108 MB
Likes
0
Public
Click a slice to open those files.
.bin54 MB · 98%
From the Hugging Face model README
This model is fine-tuned for multi-label text classification of Online Risks in Instant Messenger dialogs of Adolescents.
The model was fine-tuned on a dataset of Instant Messenger dialogs of Adolescents. The classification is multi-label and the model outputs probablities for labels {0,1,2,3,4,5}:
Here is how to use this model to classify a context-window of a dialogue:
import numpy as np
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# Prepare input texts. This model is pretrained on multi-lingual data
# and fine-tuned on English
test_texts = ['Utterance1;Utterance2;Utterance3']
# Load the model and tokenizer
model = AutoModelForSequenceClassification.from_pretrained(
'justtherightsize/small-e-czech-multi-label-online-risks-cs', num_labels=6).to("cuda")
tokenizer = AutoTokenizer.from_pretrained(
'justtherightsize/small-e-czech-multi-label-online-risks-cs',
use_fast=False, truncation_side='left')
assert tokenizer.truncation_side == 'left'
# Define helper functions
def predict_one(text: str, tok, mod, threshold=0.5):
encoding = tok(text, return_tensors="pt", truncation=True, padding=True,
max_length=256)
encoding = {k: v.to(mod.device) for k, v in encoding.items()}
outputs = mod(**encoding)
logits = outputs.logits
sigmoid = torch.nn.Sigmoid()
probs = sigmoid(logits.squeeze().cpu())
predictions = np.zeros(probs.shape)
predictions[np.where(probs >= threshold)] = 1
return predictions, probs
def print_predictions(texts):
preds = [predict_one(tt, tokenizer, model) for tt in texts]
for c, p in preds:
print(f'{c}: {p.tolist():.4f}')
# Run the prediction
print_predictions(test_texts)