Downloads · 30 days
21
11% of all-time downloads
GOSHUNCLE/ic_content_firewall_zh
ic_content_firewall_zh is a text generation model from GOSHUNCLE. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
<img src="https://huggingface.co/GOSHUNCLE/iccontentfirewallzh/resolve/main/banner.svg" alt="banner" width="900" /
Downloads · 30 days
21
11% of all-time downloads
All-time downloads
183
Public
Repo size
73.9 MB
Likes
0
Public
Click a slice to open those files.
.safetensors73.9 MB · 100%
From the Hugging Face model README
A LoRA adapter on top of Qwen2.5-1.5B-Instruct that performs structured content-moderation analysis for IC-design industry text (mixed Traditional Chinese + English). It outputs:
分析中英文混合的 IC 設計業工程師常見機密資訊用語,有效避免企業資敏資訊外洩,輸出結構化 JSON:
| 維度 | 內容 |
|---|---|
| 內容分類 | RTL / CUSTOMER / QUOTE / VENDOR / PROCESS / SCHEDULE / TESTING / INTERNAL / PUBLIC(multi-label) |
| 風險等級 | L0 公開 / L1 內部 / L2 需審核 / L3 機密 |
| 機敏實體 | CUSTOMER, PROJECT, VENDOR, PRICE, PROCESS_NODE, MODULE_NAME, IP_BLOCK, YIELD_NUMBER, SPEC_PARAM |
| 原因說明 | 翻譯式(將技術術語轉換成具體意涵,協助非技術人員快速理解) |
from inference import Detector
det = Detector() # 載入 base model + adapter
result = det.detect("Codename-Alpha 預計 5/15 RTL freeze,pcie_phy 還有 corner case 沒清。")
print(result)
# {
# "content_categories": ["RTL", "SCHEDULE"],
# "risk_level": "L3",
# "sensitive_entities": [
# {"type": "PROJECT", "value": "Codename-Alpha", "reason": "..."},
# {"type": "MODULE_NAME", "value": "pcie_phy", "reason": "..."}
# ],
# "reasoning": {
# "category_reason": "...",
# "risk_reason": "..."
# }
# }
| L1: base + no Filter | L2: adapter + no Filter | L3: adapter + Filter | |
|---|---|---|---|
| 格式正確率 | 99.0% | 100.0% | 100.0% |
| Risk Accuracy | 28.3% | 97.0% | 97.0% |
| Risk Within ±1 | 90.9% | 100.0% | 100.0% |
| Cat Macro F1 | 25.2% | 94.0% | 94.0% |
| Cat Exact Match | 5.1% | 70.7% | 70.7% |
| Entity Micro F1 | 9.0% | 95.5% | 95.5% |
| Entity Precision | 26.9% | 96.5% | 96.5% |
| Entity Recall | 5.4% | 94.6% | 94.6% |
L1 → L2 大幅躍升證明微調有效;L2 = L3 代表微調後的模型輸出已乾淨到 Filter 不需介入(但 Filter 仍為保險網保留)。
模型輸出之後再經過三層補強:
value 必須逐字出現在原文,防止幻覺A LoRA adapter producing structured content-moderation analysis for IC-design industry text (mixed Traditional Chinese + English):
from inference import Detector
det = Detector()
result = det.detect("Codename-Alpha tape-out scheduled 2026 Q3 on Vendor-X N3 process.")
print(result)
| L1: base, no Filter | L2: adapter, no Filter | L3: adapter + Filter | |
|---|---|---|---|
| Format ok | 99.0% | 100.0% | 100.0% |
| Risk Accuracy | 28.3% | 97.0% | 97.0% |
| Cat Macro F1 | 25.2% | 94.0% | 94.0% |
| Entity Micro F1 | 9.0% | 95.5% | 95.5% |
Jump from L1 to L2 demonstrates that the fine-tuning learned the task; L2 ≈ L3 means the adapter's output is already clean enough that the universal filters don't need to intervene on in-distribution data (filters retained as safety net).
Output passes through three augmentation layers:
value must appear verbatim in the input (anti-hallucination)risk_level.Apache 2.0. The model weights, inference code, and demo are all released under Apache 2.0.
Training data: All 600 training samples and 99 holdout samples were synthesized via template + slot-filling. The default demo dictionary in the companion Space uses fictional placeholder names (e.g., Customer-A, Vendor-Foundry-X, Codename-Alpha); none refer to any real company, vendor, or project. Users deploying this model should replace the demo dictionary with their own organization's actual lists. The model's outputs do not represent any specific company, vendor, or project. The authors disclaim any responsibility for misuse.