Downloads · 30 days
3
5% of all-time downloads
ishank9/lora-fraud-classifier
lora-fraud-classifier is a text classification model from ishank9. Use it when you need a label for a piece of text. It is set up for peft. The card lists the license as apache-2.0.
This model is a LoRA fine-tuned version of Qwen2.5-1.5B-Instruct, adapted to classify KYC/transaction notes as suspicious or notsuspicious, with a one-line reason — inspired by real-world fraud detection and identity…
Downloads · 30 days
3
5% of all-time downloads
All-time downloads
65
Public
Repo size
20.2 MB
Likes
1
Public
Click a slice to open those files.
.json11.4 MB · 72%
From the Hugging Face model README
This model is a LoRA fine-tuned version of Qwen2.5-1.5B-Instruct,
adapted to classify KYC/transaction notes as suspicious or not_suspicious, with a one-line reason —
inspired by real-world fraud detection and identity verification use cases.
q_proj, v_proj), r=16, alpha=32| Stage | Accuracy |
|---|---|
| Zero-shot baseline (no fine-tuning) | ~60-70% (biased toward false positives) |
| After LoRA fine-tuning | 100% (53/53 on held-out test set) |
The base model consistently over-flagged routine transactions (groceries, rent, tax refunds) as suspicious. Fine-tuning on balanced examples corrected this bias.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base_model, "ishank9/lora-fraud-classifier")
tokenizer = AutoTokenizer.from_pretrained("ishank9/lora-fraud-classifier")
prompt = "Classify the following note as 'suspicious' or 'not_suspicious' and give a one-line reason.\n\nNote: A dormant account suddenly received ₹900,000 and transferred it out the same day.\n\nAnswer:"
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(output[0], skip_special_tokens=True))
This model was trained on a relatively small, LLM-generated synthetic dataset. While it performs strongly on similarly-styled held-out examples, this reflects pattern-learning on clean synthetic data rather than a guarantee of real-world generalization to messier, real transaction data.