Downloads · 30 days
12
26% of all-time downloads
niranjan2777/SENTINEL
SENTINEL is a text generation model from niranjan2777. Use it when you need the model to write or continue text. It is set up for peft.
This repository contains the Final (QLoRA SFT + GRPO) adapter weights for SENTINEL, an autonomous web-exploitation agent.
Downloads · 30 days
12
26% of all-time downloads
All-time downloads
47
Public
Repo size
185 MB
Likes
0
Public
Click a slice to open those files.
.safetensors168 MB · 91%
From the Hugging Face model README
This repository contains the Final (QLoRA SFT + GRPO) adapter weights for SENTINEL, an autonomous web-exploitation agent.
SENTINEL is trained to autonomously navigate, analyze, and exploit web vulnerabilities using a structured JSON-based reasoning and action schema .
This repository represents the completion of both Stage 1 and Stage 2:
The model was fine-tuned using Unsloth for optimized, 2x faster training.
unsloth/llama-3-8b-Instruct-bnb-4bitLoss curve during training:
| Step | Training Loss | Validation Loss |
|---|---|---|
| 20 | 1.333500 | 1.492488 |
| 40 | 1.127200 | 1.263251 |
| 60 | 0.689200 | 1.210084 |
| 80 | 0.724800 | 1.163273 |
| 100 | 0.763000 | 1.146508 |
The model was trained on a custom dataset (train_llama3.jsonl) consisting of 415 SENTINEL trajectory pairs. These trajectories represent successful pentesting workflows, teaching the model how to target vulnerability sinks (form actions, hidden fields, query parameters, JSON bodies, etc.), infer backend technologies, and deliver appropriate payloads.
Because this is a PEFT (Parameter-Efficient Fine-Tuning) adapter, you must load the base model (unsloth/llama-3-8b-Instruct-bnb-4bit or the standard meta-llama/Meta-Llama-3-8B-Instruct) and apply these LoRA weights on top using the peft library.