Downloads · 30 days
0
danielhjerresen/BDS_M4_exam_final_model
BDS_M4_exam_final_model is a text classification model from danielhjerresen. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This repository contains the final fine-tuned PatentSBERTa model developed for the Advanced Agentic Workflow with QLoRA final project.
Downloads · 30 days
0
Access
Public
Updated Mar 4, 2026
Repo size
438 MB
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
This repository contains the final fine-tuned PatentSBERTa model developed for the Advanced Agentic Workflow with QLoRA final project.
The model classifies patent claims as green vs non-green technologies, focusing on climate mitigation technologies aligned with CPC Y02 classifications.
The training pipeline combines silver labels, agent debate labeling, and targeted human review to improve classification quality on difficult claims.
Base model: AI-Growth-Lab/PatentSBERTa
Task: Binary classification
Labels:
| Label | Meaning |
|---|---|
| 0 | Non-green technology |
| 1 | Green technology (climate mitigation related) |
The model is fine-tuned using HuggingFace AutoModelForSequenceClassification.
The training dataset is based on a balanced 50k patent claim dataset derived from:
AI-Growth-Lab/patents_claims_1.5m_train_test
Dataset composition:
| Source | Description |
|---|---|
| Silver Labels | Automatically derived from CPC Y02 indicators |
| Gold Labels | 100 high-uncertainty claims reviewed using MAS + Human-in-the-Loop |
Final training set:
train_silver + gold_100
The gold dataset overrides the silver labels for those claims to improve supervision on ambiguous cases.
The full system used in the project consists of several stages:
The model is evaluated on the eval_silver split of the dataset.
Primary metric:
F1 score
Additional metrics reported:
The evaluation script also exports prediction probabilities for analysis.
Example inference using HuggingFace Transformers:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "YOUR_USERNAME/BDS_M4_exam_final_model"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "A system for capturing carbon emissions using advanced filtration..."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
outputs = model(**inputs)
prob_green = torch.softmax(outputs.logits, dim=-1)[0,1].item()
print("Probability green:", prob_green)
This model was developed as part of the M4 Advanced AI Systems final assignment.
The project explores agentic workflows for data labeling, combining:
If referencing this model in academic work:
Green Patent Detection with Agentic Workflows.
M4 Advanced AI Systems Final Project.
Student project submission by Daniel Hjerresen.