Downloads · 30 days
0
Pranavkk/AntiAtropos
AntiAtropos is a machine learning model from Pranavkk. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A production-grade RL/agent environment for the future of autonomous DevOps — where intelligent agents replace fragile runbooks, reduce on-call toil, and keep infrastructure healthy without human intervention.
Downloads · 30 days
0
Access
Public
Updated Apr 25, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.lock621 KB · 44%
From the Hugging Face model README
A production-grade RL/agent environment for the future of autonomous DevOps — where intelligent agents replace fragile runbooks, reduce on-call toil, and keep infrastructure healthy without human intervention.
AntiAtropos is an open, high-fidelity environment for training and benchmarking AI agents on site reliability engineering (SRE) — the discipline that keeps production infrastructure alive at scale. It models a live five-node microservice cluster operating under realistic production pressures: demand surges, cascading node failures, SLA deadlines, and hard safety constraints on critical services.
This is not a toy grid world or an abstract planning problem. Every action type, every penalty function, and every telemetry field in AntiAtropos was designed to mirror the exact decisions an on-call engineer faces when the PagerDuty alert fires at 3 AM.
Modern platform teams operate infrastructure that is orders of magnitude more complex than the teams managing it. The result is a well-documented set of pain points:
AntiAtropos is a training and evaluation ground for agents that solve this problem — systems that can observe cluster telemetry, reason about multi-step consequences, and issue control actions that keep services healthy, cost-efficient, and resilient.
An agent operating in AntiAtropos executes the same core loop a platform engineer runs continuously:
1. Observe the cluster state. The observation space mirrors real Prometheus/Grafana metrics: request rates, p99 latency, error rates, queue backlogs, CPU utilization, and per-node health — the same signals that drive every serious SRE incident workflow.
2. Reason about what is wrong and why. The environment implements genuine queueing dynamics with boot delays, traffic reroute decay, and Lyapunov-based stability measurement. Agents that only react to threshold breaches perform poorly; agents that build a causal model of the cluster perform well.
3. Issue control actions with real operational semantics.
SCALE_UP — expand node capacity (with a realistic BOOT_DELAY_TICKS = 5 cold-start delay)SCALE_DOWN — reduce capacity and costREROUTE_TRAFFIC — shift request load away from unhealthy nodesSHED_LOAD — drop a fraction of traffic to protect the cluster (forbidden on critical nodes)NO_OP — hold position when the system is stableThese are not abstract symbols. They map directly to kubectl scale, traffic policy overrides, and rate limiter controls used in production Kubernetes environments.
4. Balance competing objectives across time. Uptime vs. cost vs. stability is the fundamental trade-off every platform team navigates. Brute-force overprovisioning fails the cost grader. Underprovisioning fails SLAs. The agent must plan — not just react.
5. Respect hard safety constraints. Critical nodes cannot have load shed. Scale operations are bounded. Invalid actions are penalized. AntiAtropos enforces the same guardrails that production runbooks encode, rewarding agents that understand operational boundaries.
The trajectory of platform engineering is clear: the toil layer gets automated, and engineers move up the stack. AntiAtropos provides the training and evaluation infrastructure to accelerate that transition responsibly:
control/kubernetes_executor.py) and Prometheus ingestion (telemetry/prometheus_client.py) make it possible to wire a trained policy into a real cluster with minimal adaptation, enabling true hybrid-autonomy workflows where the agent handles routine incidents and escalates novel ones.[0.0, 1.0] composite score per episode — making benchmark comparisons across models and policies reproducible and auditable.inference.py provides a complete evaluation harness for testing frontier LLMs as zero-shot SRE agents. Set your API key, pick a model, and run — the script handles the full episode loop: observation formatting, action parsing, constraint enforcement, and final grading.
set OPENAI_API_KEY=your_key_here
set MODEL_NAME=gpt-4.1
set ANTIATROPOS_TASK=task-3
python inference.py
This makes AntiAtropos a drop-in benchmark for comparing how well different LLMs reason about infrastructure operations — a capability that is increasingly relevant as AI models are integrated into on-call tooling, runbook automation, and incident triage systems.
AntiAtropos implements typed OpenEnv interfaces using Pydantic models and an OpenEnv-compatible FastAPI server:
Action model: SREAction in models.pyObservation model: ClusterObservation + NodeObservation in models.pystep(action) returns observation with reward/done fieldsreset() returns initial cluster observationstate is exposed through the OpenEnv State objectopenenv.yaml is present at repository rootOpenEnv manifest:
AntiAtroposfastapiserver.app:app7860For each node i, AntiAtropos uses a fluid queue update:
Q_i(t+1) = max(Q_i(t) + lambda_eff_i(t) - mu_i(t), 0)
where:
lambda_eff_i = lambda_incoming_i * (1 - shed_fraction_i)mu_i = capacity_i * 15 requests/tick (or 0 if node has failed)cpu_i = lambda_incoming_i / mu_ilatency_i = BASE_LATENCY_MS + LATENCY_STEEPNESS * Q_iCore stability objective is weighted Lyapunov energy:
V(s) = sum_i (w_i * Q_i^2)
VIP/business-critical nodes carry higher weights w_i. Drift term:
DeltaV(t) = V(s_t) - V(s_{t-1})
R_raw_t = -(alpha * DeltaV_t + beta * Cost_t + gamma * SLA_violation_t)
Default weights: alpha = 0.002, beta = 0.01, gamma = 10.0.
Normalized:
R_norm_t = sigmoid((R_raw_t - midpoint) / temperature)
Dense step-level signal — not sparse terminal reward — that strongly penalizes SLA failures, invalid actions, and destabilizing queue growth.
SREAction (models.py):
action_type: NO_OP | SCALE_UP | SCALE_DOWN | REROUTE_TRAFFIC | SHED_LOADtarget_node_id: node-0 to node-4parameter: bounded float with action-dependent semanticsSafety constraints enforced by control/validation.py:
SHED_LOAD is forbidden on critical nodes (node-0, node-1, node-2)ClusterObservation:
NodeObservation listNodeObservation (per node):
task-1 — Capacity Ramp (Easy)Load starts near cluster capacity and ramps over the episode. The agent must proactively scale and contain queue growth without overprovisioning. A clean benchmark for predictive capacity planning — the most common form of infrastructure toil.
task-2 — Fault Tolerance (Medium)A non-VIP node fails at a randomized tick. Traffic continues hitting failed capacity until the agent detects the failure and responds. Tests reactive incident response: detecting failure signals, rerouting affected traffic, and compensating with scaling — under realistic delay constraints.
task-3 — Stability Under Surge (Hard)Major traffic surges target non-critical nodes, threatening to cascade. The agent must protect the VIP Payment Gateway (node-0). SHED_LOAD is forbidden on critical nodes (node-0, node-1, and node-2). The agent must coordinate pre-emptive SCALE_UP to absorb the surge before it arrives and use persistent REROUTE_TRAFFIC to redirect load, all while maintaining cost discipline. The closest analogue to a real high-severity incident: time pressure, safety constraints, and no single correct action.
Computed by grader.py — deterministic and reproducible:
| Component | Formula | Weight |
|---|---|---|
| Uptime | Fraction of steps with latency ≤ 0.20 and error rate ≤ 0.05 | 0.4 |
| Cost | exp(-3.0 * over_ratio) — punishes overprovisioning | 0.4 |
| Stability | 1 / (1 + (avg_energy / TARGET_ENERGY)^power) | 0.2 |
< 0.5-0.05 per invalid action0.0AntiAtropos ships a full production-style observability stack:
GET /metricsantiatropos-overview dashboard: reward trajectory, queue heatmaps, latency timeseries, SLA violations, per-node state, action throughput, executor reliability/, /prometheus/, and /grafana/ on port 7860deploy/entrypoint.sh boots the full stack in a single containerFor teams evaluating agents against real infrastructure:
control/kubernetes_executor.py translates SCALE_UP/SCALE_DOWN into kubectl operations on mapped deploymentsANTIATROPOS_WORKLOAD_MAP or ANTIATROPOS_NODE_DEPLOYMENT_MAPtelemetry/prometheus_client.py ingests live PromQL metrics and reconciles them into simulator state via weighted blending — enabling a real-environment feedback loop with minimal code changeReproducible NO-OP baseline over 20 seeded runs (100 steps each):
| Task | Mean Composite | Min | Max |
|---|---|---|---|
| task-1 | 0.6980 | 0.6845 | 0.7171 |
| task-2 | 0.7020 | 0.6400 | 0.7560 |
| task-3 | 0.2063 | 0.1721 | 0.2521 |
Task-3's low baseline score reflects the genuine difficulty of burst surge management under safety constraints — and the substantial headroom available for capable agents.
pip install -e .
uvicorn server.app:app --host 0.0.0.0 --port 8000
docker build -t antiatropos:latest .
docker run --rm -p 7860:7860 antiatropos:latest
openenv validate
openenv push
| Path | Description |
|---|---|
models.py | Typed OpenEnv action/observation models |
simulator.py | Queueing physics, task dynamics, action semantics |
stability.py | Lyapunov/reward math |
grader.py | Deterministic episode scoring |
inference.py | OpenAI-compatible baseline runner |
client.py | OpenEnv client wrapper |
openenv.yaml | Environment manifest |
server/AntiAtropos_environment.py | Environment runtime (reset, step, state handling) |
server/app.py | FastAPI/OpenEnv app + /metrics |
control/ | Action validation and Kubernetes executor |
telemetry/ | Prometheus ingestion, metric mapping, exporter |
deploy/ | Entrypoint, NGINX, Prometheus, Grafana provisioning |
Dockerfile, server/Dockerfile | Container build targets |
Dockerfile, server/Dockerfile)uv.lock)grader.py)For fixed-seed studies, use controlled simulator seeding in evaluation harnesses.
| Criterion | AntiAtropos |
|---|---|
| Real-world utility | Genuine SRE/platform engineering control task with production-grade operational constraints |
| Task quality | 3 tasks with easy-medium-hard progression mapped to real incident categories |
| Grader quality | Deterministic, interpretable composite score in [0, 1] |
| Environment design | Dense Lyapunov-grounded reward, clean reset/step loop, explicit episode boundaries |
| Code quality | Typed Pydantic models, modular components, OpenEnv manifest, containerized runtime |
| Novelty | Lyapunov reward shaping + live K8s control plane + Prometheus telemetry + observability-first design |