Downloads · 30 days
0
Yaxh07/rl-cyber
rl-cyber is a machine learning model from Yaxh07. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Research-grade OpenEnv environment for autonomous cyber defense using reinforcement learning. This project models SOC workflows (alert triage, containment, patching) as a sequential decision process with attacker-defe…
Downloads · 30 days
0
Access
Public
Updated Apr 8, 2026
Repo size
11.8 MB
Likes
0
Public
Click a slice to open those files.
.zip12.7 MB · 99%
From the Hugging Face model README
Research-grade OpenEnv environment for autonomous cyber defense using reinforcement learning. This project models SOC workflows (alert triage, containment, patching) as a sequential decision process with attacker-defender interaction.
CyberSimulator (core): deterministic cyber world transitions and reward shaping.CyberEnvironment (OpenEnv server): exposes reset, step, state endpoints.DefenderGymEnv / AttackerGymEnv: RL wrappers for PPO/DQN training.TaskGrader: deterministic score in [0.0, 1.0] with transparent decomposition.CyberObservation)host_compromise: host -> compromised boolhost_isolation: host -> isolated boolservice_status: host -> service patched/healthy boolids_alerts: simulated SIEM/IDS alert stringstraffic_anomaly_score: float [0, 1]active_incidents: typed incident listreward_signal: typed reward decomposition (RewardSignal)step_budget_remaining: remaining episode budgetCyberAction)actor: defender or attackeraction_type:
block_ip, isolate_node, scan_host, patch_service, ignorelateral_move, credential_stuff, malware_drop, recon, idletarget_host, optional target_service, optional source_ipCyberState)seed=101): single-host intrusion triage and containment.seed=202): lateral movement in subnet with patch/isolation tradeoffs.seed=303): multi-stage campaign with decoy signals and constrained budget.Each task has:
[0.0, 1.0].Reward includes:
This provides feedback throughout trajectories, not only terminal reward.
Reward is typed and auditable via RewardSignal in both observation and step info.
cd C:\Users\yashm\Documents\Playground\cyber-openenv-rl
python -m pip install -e .[dev]
openenv validate .
python -m server.app --port 8000
openenv validate --url http://localhost:8000
python -m cyber_openenv_rl.rl.train_defender --algorithm ppo --task medium --timesteps 10000 --device cuda
python -m cyber_openenv_rl.rl.train_defender --algorithm dqn --task medium --timesteps 10000 --device cuda
python -m cyber_openenv_rl.rl.train_attacker --task hard --timesteps 12000
python -m cyber_openenv_rl.rl.train_curriculum --algorithm ppo --timesteps-per-task 8000 --device cuda
Training outputs include:
outputs/models/**/checkpoints)outputs/models/**/best)outputs/models/**/eval_logs)summary_*.json)Run multi-seed benchmark suite (deterministic evaluation) and generate JSON + Markdown reports:
python -m cyber_openenv_rl.eval.benchmark_suite --algorithms ppo,dqn --seeds 42,1337,2026 --timesteps 3000 --output outputs/evals/benchmark_results.json --train
Artifacts:
outputs/evals/benchmark_results.jsonoutputs/evals/benchmark_report.mdReads credentials from HF_TOKEN (required by hackathon spec).
set HF_TOKEN=your_api_key_here
python -m cyber_openenv_rl.eval.baseline_inference --model gpt-4o-mini
Output is written to outputs/evals/baseline_scores.json with per-task and aggregate scores.
The baseline uses fixed task seeds and deterministic request settings (temperature=0, top_p=1, request seed).
This project is designed for real-world defense deployment in stages:
Yes. Training is done in the simulated cyber environment (CyberSimulator + Gym wrappers), which is required for safe RL exploration.
Yes, for defensive workflows. Use the real-time deployment path:
Current implementation supports real-time defensive recommendations with confidence scores and guardrail enforcement. You can integrate your own execution connector for production remediation actions.
python -m cyber_openenv_rl.rl.train_curriculum --algorithm ppo --timesteps-per-task 15000 --device cuda
Uses NSL-KDD intrusion dataset to calibrate simulator threat profile, then trains for wall-clock duration.
python -m cyber_openenv_rl.rl.train_real_data --algorithm ppo --task hard --hours 3 --chunk-timesteps 12000 --seed 42 --device cuda --output-dir outputs/models/real_data
If CUDA is unavailable in your PyTorch install, the script will raise an error so you can fix your CUDA setup first.
python -m cyber_openenv_rl.deployment.run_realtime_defense ^
--model-path outputs/models/curriculum/hard/defender_ppo_hard.zip ^
--algorithm ppo ^
--task hard ^
--telemetry data/sample_telemetry/incident_001.json ^
--output outputs/evals/realtime_inference.json ^
--confidence-threshold 0.55
Guardrails are configured in:
configs/production_policy.yamlGuardrails block non-defensive/offensive behavior and can require approval for disruptive actions (block_ip, isolate_node).
To regenerate the project-wide index:
python tools/generate_codebase_index.py
Generated file:
CODEBASE_INDEX.mdGenerated from fixed-seed scripted defender (python -m cyber_openenv_rl.eval.scripted_baseline):
| Task | Score |
|---|---|
| easy | 0.5050 |
| medium | 0.3350 |
| hard | 0.5729 |
| aggregate | 0.4710 |
pytest -q
Test coverage includes:
openenv validate command pass.docker build -t cyber-openenv-rl -f server/Dockerfile .
docker run --rm -p 8000:8000 cyber-openenv-rl
Tag your Space with openenv and use openenv push for deployment.