Downloads · 30 days
11
22% of all-time downloads
kwoncho/ko-sroberta-korean-time-expression-classifier
ko-sroberta-korean-time-expression-classifier is a token classification model from kwoncho. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as other.
This model detects Korean TIMEX3 time expressions with BIO token classification labels.
Downloads · 30 days
11
22% of all-time downloads
All-time downloads
49
Public
Parameters
110M
440 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors440 MB · 99%
From the Hugging Face model README
This model detects Korean TIMEX3 time expressions with BIO token classification labels.
The backbone is jhgan/ko-sroberta-multitask, fine-tuned on 158.시간 표현 탐지 데이터 for four TIMEX3 entity types:
DATETIMEDURATIONSETUse this model to identify Korean time expressions in sentences or utterances. It predicts token-level BIO labels and can be used through the Hugging Face token-classification pipeline.
This is an experimental model trained for TIMEX3 span detection. It does not extract EVENT or TLINK annotations.
The model was trained on the official Training split and evaluated on the official Validation split of 158.시간 표현 탐지 데이터.
Training/evaluation preprocessing:
max_length=256 are excluded.text fields are used as the source text.python -m time_expression_classifier.train_token_classifier \
--data-root "158.시간 표현 탐지 데이터" \
--model-name jhgan/ko-sroberta-multitask \
--output-dir outputs/official_epoch2 \
--split-mode official \
--epochs 2 \
--learning-rate 3e-5 \
--batch-size 16 \
--max-length 256
Key settings:
| setting | value |
|---|---|
| backbone | jhgan/ko-sroberta-multitask |
| epochs | 2 |
| learning rate | 3e-5 |
| batch size | 16 |
| max length | 256 |
| weight decay | 0.01 |
| warmup ratio | 0.06 |
| seed | 42 |
Metrics are entity-level exact match on the official Validation split.
| metric | value |
|---|---|
| entity precision | 0.8265 |
| entity recall | 0.8268 |
| entity F1 | 0.8266 |
| token accuracy | 0.9899 |
| eval loss | 0.0350 |
Per-label entity-level results:
| label | precision | recall | F1 | support |
|---|---|---|---|---|
| DATE | 0.8495 | 0.8367 | 0.8430 | 23422 |
| TIME | 0.7933 | 0.8033 | 0.7983 | 3665 |
| DURATION | 0.7848 | 0.8247 | 0.8042 | 6810 |
| SET | 0.7107 | 0.6910 | 0.7007 | 974 |
from transformers import pipeline
tagger = pipeline(
"token-classification",
model="kwoncho/ko-sroberta-korean-time-expression-classifier",
aggregation_strategy="simple",
)
text = "매주 토요일 저녁에 회의를 합니다."
print(tagger(text))
주, 하루, 시간, 한달, 일주일, and 매일.SET is the lowest-performing label due to smaller support and ambiguity between repeated events and duration expressions.Repository: [email protected]:hyun2019/ko-sroberta-korean-time-expression-classifier.git
The local release artifact is tracked as models/official_epoch2 via DVC.