Monitoring, logging, alerting, and site reliability engineering reference.
Observability reference: Prometheus for metrics, Loki for logs, Grafana for visualization, Mimir for long-term storage, plus architecture tradeoffs and deployment patterns.
Site reliability engineering patterns: monolithic vs microservice tradeoffs, deployment strategies, SLOs and error budgets, multi-cluster architectures, and incident management workflows.
OpenTelemetry reference: SDK setup, collector architecture, auto-instrumentation patterns, sampling strategies, and integration with the Grafana LGTM stack.