Downloads · 30 days
290
2% of all-time downloads
liwii/fluency-score-classification-ja
fluency-score-classification-ja is a machine learning model from liwii. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of line-corporation/line-distilbert-base-japanese on the "日本語文法誤りデータセット". It achieves the following results on the evaluation set: - Loss: 0.1912 - ROC AUC: 0.9811
Downloads · 30 days
290
2% of all-time downloads
All-time downloads
11.6K
Public
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.bin275 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of line-corporation/line-distilbert-base-japanese on the "日本語文法誤りデータセット". It achieves the following results on the evaluation set:
This model wraps line-corporation/line-distilbert-base-japanese with DistilBertForSequenceClassification to make a binary classifier.
This model can be used to classify whether the given Japanese texts are fluent (i.e., not having grammactical errors). Example usage:
# Load the tokenizer & the model
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("line-corporation/line-distilbert-base-japanese", trust_remote_code=True)
model = AutoModelForSequenceClassification.from_pretrained("liwii/fluency-score-classification-ja")
# Make predictions
input_tokens = tokenizer([
'黒い猫が',
'黒い猫がいます',
'あっちの方で黒い猫があくびをしています',
'あっちの方でで黒い猫ががあくびをしています',
'ある日の暮方の事である。一人の下人が、羅生門の下で雨やみを待っていた。'
],
return_tensors='pt',
padding=True)
output = model(**input_tokens)
with torch.no_grad():
# Probabilities of [not_fluent, fluent]
probs = torch.nn.functional.softmax(
output.logits, dim=1)
probs[:, 1] # => tensor([0.1007, 0.2416, 0.5635, 0.0453, 0.7701])
The scores could be low for short sentences even if they do not contain any grammatical erros because the training dataset consist of long sentences.
From "日本語文法誤りデータセット", used 512 rows as the evaluation dataset and the rest of the dataset as the training dataset. For each dataset split, Used the "original" rows as the data with "fluent" label, and "perturbed" as the data with "not fluent" data.
Fine-tuned the model for 5 epochs. Freezed the params in the original DistilBERT during the fine-duning.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Roc Auc |
|---|---|---|---|---|
| 0.4582 | 1.0 | 647 | 0.2887 | 0.9679 |
| 0.2664 | 2.0 | 1294 | 0.2224 | 0.9761 |
| 0.2177 | 3.0 | 1941 | 0.2047 | 0.9793 |
| 0.1899 | 4.0 | 2588 | 0.1944 | 0.9807 |
| 0.1865 | 5.0 | 3235 | 0.1912 | 0.9811 |