Monitoring & Observability
monitoring-observability
Monitoring and observability patterns. Prometheus RED/USE metrics, structured logging with Pino/Winston, OpenTelemetry tracing, SLO-based alerting, Grafana dashboards, burn rate alerts.
SKILL.md
Full skill instructions
Monitoring & Observability
When to use
Use when setting up metrics, logging, tracing, or alerting for services. Covers the three pillars of observability, SLO definition, dashboard design, and alert fatigue prevention.
Core principles
- Three pillars: logs, metrics, traces — each answers different questions
- Structured logging always — JSON logs are searchable; free-text logs are archaeology
- RED method for services — Rate, Errors, Duration per endpoint
- SLO-based alerting — alert on burn rate, not raw error count
- Correlation IDs through every hop — without them, distributed debugging is guesswork
References available
references/metrics-patterns.md— RED/USE methods, Prometheus counters/histograms/gauges, Grafana dashboardsreferences/logging-patterns.md— Structured logging (Pino/Winston), levels, aggregation (ELK/Loki), correlation IDsreferences/tracing-patterns.md— OpenTelemetry setup, span design, sampling strategies, trace analysisreferences/alerting-strategies.md— SLO-based alerting, burn rate (14.4x fast, 6x slow), alert fatigue prevention
Assets available
assets/grafana-dashboard-template.json— Starter RED metrics dashboardassets/alert-rules-template.yaml— Prometheus alert rules for SLO burn rate
