Downloads · 30 days
12
14% of all-time downloads
Saniter/BBS-NewWordFind
BBS-NewWordFind is a text classification model from Saniter. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
基于 Qwen3.5-0.8B-Base 微调的中文论坛新词发现分类模型。用于判定语境中的候选词是否具备独立、稳定的语义,过滤残片(Fragment)、乱码(Garbled)与高频虚词(Too Common)。
Downloads · 30 days
12
14% of all-time downloads
All-time downloads
87
Public
Parameters
752M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.5 GB · 98%
From the Hugging Face model README
基于 Qwen3.5-0.8B-Base 微调的中文论坛新词发现分类模型。用于判定语境中的候选词是否具备独立、稳定的语义,过滤残片(Fragment)、乱码(Garbled)与高频虚词(Too Common)。
本模型采用定制化的特征池化策略与损失函数:
测试集指标:
推荐使用 uv 进行依赖管理与执行。
uv venv
source .venv/bin/activate
uv add torch transformers huggingface_hub loguru tqdm scikit-learn accelerate
下载仓库内的 inference.py 脚本,执行批量推理:
uv run inference.py --input data.json
脚本默认指向本仓库拉取权重(包含定制的 classifier_head.pt),并在当前目录输出 accepted_words.json 与 rejected_words.json。
data.json)输入必须包含候选词及其关联语境列表:
{
"io_uring": {
"contexts": [
"Linux 6.12 对 io_uring 做了深度优化。",
"推荐使用 io_uring 替代传统的 epoll。"
]
}
}