Downloads · 30 days
12
11% of all-time downloads
dpevzner/CyberSecurity_Microsoft_Phi3B
CyberSecurity_Microsoft_Phi3B is a text generation model from dpevzner. Use it when you need the model to write or continue text. It is set up for peft.
Cybersecurirty Assistant is a LoRA fine-tuned adapter for microsoft/Phi-3.5-mini-instruct (3.8B parameters), purpose-built for cybersecurity reasoning tasks as a proof of concept, to test my dataset to model pipeline.…
Downloads · 30 days
12
11% of all-time downloads
All-time downloads
105
Public
Repo size
17.8 MB
Likes
0
Public
Click a slice to open those files.
.safetensors17.8 MB · 83%
From the Hugging Face model README
Cybersecurirty Assistant is a LoRA fine-tuned adapter for microsoft/Phi-3.5-mini-instruct (3.8B parameters), purpose-built for cybersecurity reasoning tasks as a proof of concept, to test my dataset to model pipeline. The model is trained to assist with red team tradecraft analysis, blue team detection reasoning, PowerShell/CMD/Bash command interpretation, network analysis (nmap, Wireshark, tcpdump), and Sigma rule evaluation.
The adapter targets a 3B-class parameter range using QLoRA 4-bit quantization, making it practical to run on consumer-grade hardware with 8GB VRAM. It is the first completed training run in a sequenced multi-model project that plans to also train Mistral 7B Instruct v0.3, Llama 3.1 8B Instruct, Qwen2.5 7B Instruct, DeepSeek-R1-Distill-Llama-8B, and Mistral NeMo 12B Instruct on the same corpus.
train_lora.py, dashboard.ps1 (internal lab)Cybersecurirty Assistant is designed for use in a controlled cybersecurity lab environment. It is intended to assist a trained security professional with:
The adapter is designed to be merged into its base model using PEFT merge_and_unload and served via Ollama + Open WebUI for network-accessible inference in a homelab environment. Downstream integration into an agentic scaffold with a human-in-the-loop validation layer is the intended deployment path.
Users must be security professionals operating within authorized lab or engagement boundaries. The model requires a human validation layer for any tool execution suggestions. Do not use for autonomous operations or outside approved scope.
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch
base_model_id = "microsoft/Phi-3.5-mini-instruct"
adapter_path = "./adapter" # path to saved LoRA adapter weights
tokenizer = AutoTokenizer.from_pretrained(base_model_id)
model = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
model = PeftModel.from_pretrained(model, adapter_path)
# Example prompt
prompt = "What does `Get-NetTCPConnection | Where-Object State -eq 'Established'` return, and how should a SOC analyst interpret the output?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(output[0], skip_special_tokens=True))
The training dataset is a synthetic cybersecurity corpus in JSONL format, organized into three epoch families:
| Epoch | Folder | Contents | Records |
|---|---|---|---|
| 1 | 03_toolknowledge + 05_detection_rules | Tool syntax references, Sigma/YARA detection rules, tradecraft matrix, subnetting, tcpdump/Wireshark reasoning | 97 |
| 2 | 04_goldens | Execution traces — success and failure cases for Linux privesc, Windows initial access, nmap, persistence, network analysis, evasion | 87 |
| 3 | 06_contrast_pairs | Ambiguity and intent reasoning — admin vs. malicious command pairs, Linux/Windows privesc contrast, cross-platform equivalence | 23 |
| Total | 207 |
Domain coverage at time of Run 1:
| Domain | Records | Target | Coverage |
|---|---|---|---|
| PowerShell Administration | ~50 | 550 | 9% |
| Red Team & Offensive Tradecraft | ~50 | 220 | 23% |
| Linux / POSIX Operations | ~60 | 150 | 40% |
| Windows Command Line | ~45 | 85 | 53% |
| Network Analysis (nmap, tcpdump, Wireshark) | ~30 | 150 | 20% |
| Cybersecurity Reasoning / Ambiguity | ~60 | 200 | 30% |
| Sigma Detection Rules | 35 | 200 | 18% |
| Blue Team / SOC Workflows | ~40 | 200 | 20% |
| IPv4/IPv6 / Port / Service Identification | ~10 | 150 | 7% |
Source material includes extracted content from technical references covering PowerShell, Windows CMD, Linux/Bash, network analysis (Nmap, Wireshark, tcpdump), red team field manuals (RTFM), NIST cybersecurity frameworks, and Sigma detection rule libraries.
Benchmark data (07_benchmarks/) was explicitly excluded from training and held as a frozen evaluation set.
Training uses Supervised Fine-Tuning (SFT) via TRL's SFTTrainer with QLoRA (4-bit NF4 quantization).
Curriculum structure:
| Parameter | Value |
|---|---|
| Training regime | bf16 mixed precision |
| LoRA rank (r) | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| Trainable parameters | 8,912,896 / 2,018,053,120 (0.44%) |
| Max sequence length | 1024 tokens |
| Epochs | 3 |
| Total training steps | 69 |
| Formatted / ingested examples | 11 / 207 |
| Warmup ratio | configured (deprecated in v5.2 — migrated to warmup_steps) |
| Component | Specification |
|---|---|
| Machine | Alienware M16 R2 |
| CPU | Intel Core Ultra 7 155H — 16 cores / 22 logical |
| GPU | NVIDIA GeForce RTX 4070 Laptop GPU — 8GB GDDR6 VRAM |
| RAM | 64GB system memory |
| OS | Windows 11 |
| Training time | ~16 minutes (948.5 seconds) |
| Peak VRAM usage | ~4.5GB |
| Peak GPU temp | ~74°C |
Benchmark evaluation uses a frozen evaluation set stored in 07_benchmarks/frozen_eval/. This dataset was never included in training data. Benchmark families cover:
| Metric | Value |
|---|---|
| Epoch 1 loss | 8.1072 |
| Epoch 2 loss | ~6.5 |
| Epoch 3 loss | 5.2204 |
| Mean train loss | 7.17 |
| Mean token accuracy | 0.3049 |
| Loss reduction | 54.5% (11.49 → 5.22) |
| Benchmark score | 0/10 — evaluator schema mismatch (not reflective of model capability) |
Note: The benchmark evaluator in Run 1 failed to match benchmark records due to a schema format mismatch. Benchmark results from Run 1 are invalid. The benchmark handler was corrected for Run 6 onward.
| Run | Records | Epochs | Final Loss | Benchmark | Key Finding |
|---|---|---|---|---|---|
| Runs 1–3 | 207 | 3 | 5.0–6.2 | 0/10 (evaluator error) | DynamicCache bug; only 11/207 records formatted |
| Run 4 | 502 | 5 | 2.32 | 7/10 (incoherent) | Metadata contamination in V2 records |
| Run 5 | 479 | 5 | 4.18 | 7/10 (partial) | Format oscillation V1 vs V2 style |
| Run 6 | 479 | 5 | TBD | TBD | Format unified — benchmark handlers added |
Project deployment threshold: ≥85% on Failure Diagnosis AND Ambiguity/Intent Reasoning benchmarks across a dataset of 3,000+ formatted records.
Current realistic competence levels by dataset size:
| Formatted records | Expected benchmark | Agent readiness |
|---|---|---|
| ~11 (Run 1) | ~20–35% | Not deployable |
| 500 | ~45–60% | Experimental only |
| 1,500 | ~65–75% | Limited deployment |
| 3,000+ | ~80–90% | Approaching target |
| 5,000+ | 85%+ sustained | Project threshold met |
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
attn_implementation='eager' used)dashboard.ps1), per-step JSONL loss logs, hardware telemetry via pynvml + psutil