Observability isn't a dashboard you build once. It's the set of signals that tell you something is wrong before users do, and the tools that let you understand why. This category covers practical observability for production Kubernetes systems — from choosing between ELK and Loki as your logging stack, to fixing liveness and readiness probe failures that cause silent traffic loss, to building runbooks that actually get used during incidents.
Troubleshooting
Kubernetes probe failures that cause restarts or silent traffic loss.
Targets stuck DOWN — context deadline exceeded, connection refused, scheme/port mismatches, relabeling drops, and ServiceMonitor selectors.
Stop Kubernetes from killing healthy containers — covers initialDelaySeconds, timeoutSeconds, wrong paths, and startupProbe configuration.
Fix pods stuck at 0/1 Ready — covers dependency checks, probe timeouts, Istio sidecar issues, and stalled rolling deployments.
Guides
Architecture decisions and implementation guides for monitoring production systems.
Compare ELK Stack and Grafana Loki for DevOps logging — architecture, resource cost, query power, and Kubernetes fit.
Build runbooks that actually get used during incidents — covering scope, ownership, step format, and what to leave out.
All Categories
Every article on DevOps Compass is organized into a focused category.
Pod lifecycle, workloads, probes, scheduling, and production cluster operations.
EKS, IAM, ECR, VPC architecture, and cost-aware infrastructure decisions.
Jenkins, GitLab CI, GitHub Actions, deployment strategies, and automation.
Docker, image security, registries, and container runtime debugging.
Prometheus, Grafana, Loki, alerting, and observability for production.
VPCs, IAM, TLS, network policies, and access control patterns.