Downloads · 30 days
14
1% of all-time downloads
ttt421/modernbert-ja-location-classifier
modernbert-ja-location-classifier is a text classification model from ttt421. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
このモデルは、日本語の緊急通報テキストから場所タイプを多ラベル分類するために、sbintuitions/modernbert-ja-310mをファインチューニングしたものです。
Downloads · 30 days
14
1% of all-time downloads
All-time downloads
2.7K
Public
Parameters
315M
3.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt2.5 GB · 67%
From the Hugging Face model README
このモデルは、日本語の緊急通報テキストから場所タイプを多ラベル分類するために、sbintuitions/modernbert-ja-310mをファインチューニングしたものです。
CLASS_WEIGHTS = [
1.0, # apartment (236件)
1.72, # outdoor (137件)
18.15, # highway (13件)
9.08, # station (26件)
2.03, # commercial_facility (116件)
]
テストデータでの評価結果:
| クラス | Precision | Recall | F1-Score |
|---|---|---|---|
| apartment | 0.88 | 0.95 | 0.91 |
| outdoor | 0.88 | 0.83 | 0.86 |
| highway | 0.67 | 1.00 | 0.80 |
| station | 1.00 | 0.75 | 0.86 |
| commercial_facility | 0.83 | 0.71 | 0.76 |
総合スコア:
pip install transformers torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# モデルとトークナイザーのロード
model_name = "ttt421/modernbert-ja-location-classifier"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# 推論
text = "マンションの3階から火が出ています"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=1024)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.sigmoid(outputs.logits)[0]
# 結果の表示
labels = ["apartment", "outdoor", "highway", "station", "commercial_facility"]
threshold = 0.5
print("検出された場所タイプ:")
for label, prob in zip(labels, probs):
if prob > threshold:
print(f" {label}: {prob:.3f}")
texts = [
"高速道路で事故が発生しました",
"駅のホームで人が倒れています",
"ショッピングモールで迷子になりました"
]
inputs = tokenizer(texts, return_tensors="pt", truncation=True, max_length=1024, padding=True)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.sigmoid(outputs.logits)
for i, text in enumerate(texts):
print(f"
テキスト: {text}")
print("場所タイプ:")
for label, prob in zip(labels, probs[i]):
if prob > threshold:
print(f" {label}: {prob:.3f}")
highwayはテストサンプルが4件と少ないため、精度が不安定commercial_facilityのRecallが0.71と改善の余地ありApache 2.0
ベースモデル: sbintuitions/modernbert-ja-310m