Downloads · 30 days
0
LIJINGHAI111/KUAKE-LR-Rerank
KUAKE-LR-Rerank is a text classification model from LIJINGHAI111. Use it when you need a label for a piece of text. It is set up for sklearn. The card lists the license as apache-2.0.
轻量级医学文本 Rerank 模型,基于 sklearn LogisticRegression 实现 Pointwise Rerank,在 KUAKE-QTR(查询-Title 相关性)和 KUAKE-IR(医学段落检索)两个 CBLUE 基准任务上验证。
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.npz1.5 GB · 100%
From the Hugging Face model README
轻量级医学文本 Rerank 模型,基于 sklearn LogisticRegression 实现 Pointwise Rerank,在 KUAKE-QTR(查询-Title 相关性)和 KUAKE-IR(医学段落检索)两个 CBLUE 基准任务上验证。
A lightweight medical text Rerank model built with sklearn LogisticRegression (Pointwise Rerank), evaluated on the CBLUE benchmark's KUAKE-QTR (query-title relevance) and KUAKE-IR (medical passage retrieval) tasks.
| 特征组 / Group | 说明 / Description | 维度 / Dims |
|---|---|---|
| 手工特征 / Manual | overlap_ratio, jaccard, BM25, edit_dist, LCS, len_ratio 等 / etc. | 12 |
| Query TF-IDF | char_wb, ngram(1,3), sublinear_tf | 2000 |
| Title TF-IDF | 同上 / same | 2000 |
| Concat TF-IDF | query+title 拼接后 / concatenated | 4000 |
| 总计 / Total | ~8000 |
采用 TF-IDF 初筛(top-200)→ LR Rerank 两阶段流水线,全量统计评估 / Two-stage pipeline (TF-IDF top-200 → LR Rerank), full-collection evaluation:
| 指标 / Metric | 初筛 (TF-IDF) / First pass | Rerank (LR) |
|---|---|---|
| NDCG@10 | 基准 / baseline | 提升 / improved |
| NDCG@200 | — | 全量 / full |
| MRR | 基准 / baseline | 提升 / improved |
| Recall@10 | 基准 / baseline | 提升 / improved |
(具体数值请运行 infer_ir.py 查看输出 / Run infer_ir.py for exact values)
| 文件 / File | 说明 / Description |
|---|---|
lr_rerank.pkl | 完整 LR Rerank Pipeline(~1.7 GB,LFS)/ full LR rerank pipeline |
qtr_train.pkl / qtr_dev.pkl | QTR 预处理缓存 / QTR preprocessing cache |
ir_cache/ir_tfidf_vectorizer.pkl | IR 初筛 TF-IDF 向量器 / first-pass TF-IDF vectorizer |
ir_cache/ir_tfidf_matrix.npz | IR 初筛 TF-IDF 稀疏矩阵 / first-pass TF-IDF sparse matrix |
import joblib
import numpy as np
# 加载模型 / load model
pipe = joblib.load("lr_rerank.pkl")
# 输入 / inputs
queries = ["上呼吸道感染吃什么药"]
titles = ["上呼吸道感染的治疗药物"]
# 预测相关性概率 / predict relevance prob (label >= 2)
probs = pipe.predict_proba(list(zip(queries, titles)))
score = probs[:, 2] + probs[:, 3] # P(label >= 2)
print(f"相关性分数 / Relevance score: {score[0]:.4f}")
完整代码、数据集、报告位于 / Full code, data, and reports at: https://github.com/githuBlijingai/KUAKE-LR-Rerank
Apache 2.0