Step-by-step troubleshooting guides, clear tutorials, and real-world DevOps practices for Kubernetes, AWS, CI/CD, and cloud infrastructure. Less noise. More clarity.
Looking for a specific issue? Search DevOps Compass →
New here?
DevOps has a wide surface area. Whether you're new to the field or a developer who's just taken on infrastructure responsibilities, the best way to build real competence is to follow a practical path — not a theoretical one.
These four resources give you a strong foundation. Read them in order or jump straight to the one that solves your most pressing problem.
Topics
Every article is organized into a focused category so you can go deep on the stack areas that matter to your role right now.
Pod lifecycle, workloads, networking, RBAC, debugging CrashLoopBackOff, Helm, and cluster operations in production.
EKS, IAM, VPC architecture, S3, RDS, CloudWatch, Transit Gateway, and cost-aware infrastructure decisions.
Pipeline design, GitLab CI, GitHub Actions, Jenkins, deployment strategies, rollback patterns, and automation best practices.
Dockerfile optimisation, image security, multi-stage builds, Docker Compose, container registries, and runtime debugging.
Prometheus, Grafana, Loki, alerting strategy, log aggregation, SLOs, and building useful dashboards — not just busy ones.
VPCs, subnets, security groups, TLS, zero-trust principles, network policies in Kubernetes, and IAM least-privilege patterns.
Troubleshooting
Concise, structured guides written around real error messages and production failure scenarios — not textbook theory.
Pods stuck at Init:0/1 or Init:CrashLoopBackOff — failing init containers, missing ConfigMaps/Secrets, volume waits, and dependency checks.
Work through every S3 authorization layer — IAM ARNs, bucket policy, Block Public Access, Object Ownership, and KMS key permissions.
Workflows that never run — on: event and branch/path filters, invalid YAML, the GITHUB_TOKEN recursion rule, and fork PR limits.
Reclaim /var/lib/docker — prune images, containers, volumes and build cache, cap container logs, and fix inode exhaustion.
Targets stuck DOWN — context deadline exceeded, connection refused, scheme/port mismatches, relabeling drops, and ServiceMonitor selectors.
Diagnose and resolve image pull failures caused by registry credentials, wrong image tags, or network restrictions.
Step-by-step approach to trace why a container keeps crashing — from misconfigured env vars to missing secrets and bad entrypoints.
Resolve Docker socket permission errors on Jenkins agents — including the correct way to handle user groups and rootless Docker.
Fix expired tokens, IAM permission errors, cross-account access issues, and misconfigured credential helpers for AWS ECR.
Diagnose pods stuck in Pending — covering insufficient resources, taints, node selectors, affinity rules, and PVC binding failures.
Diagnose and fix Out of Memory kills — covers memory limits, JVM heap sizing, memory leaks, sidecar pressure, and right-sizing with VPA.
Stop Kubernetes from killing healthy containers — covers initialDelaySeconds, timeoutSeconds, wrong paths, health endpoint design, and startupProbe configuration.
Fix pods stuck at 0/1 Ready and removed from Service traffic — covers dependency checks, probe timeouts, Istio sidecar issues, and stalled rolling deployments.
Diagnose and fix pods stuck in Terminating — covers finalizers, force deletion, unreachable nodes, and graceful shutdown issues.
Resolve pods stuck in ContainerCreating — covers missing secrets, PVC mount failures, CNI plugin errors, and container runtime issues.
Fix AWS AccessDenied errors — covers missing IAM actions, IRSA misconfiguration, resource policies, SCPs, and the ECR token pitfall.
Fix EKS worker nodes that won't join — covers aws-auth ConfigMap, missing IAM policies, bootstrap script errors, and network connectivity.
Diagnose and fix security group rules causing silent connection timeouts — VPC flow logs, NACLs, EKS rules, and Reachability Analyzer.
Fix pods that can't reach external services — NAT Gateway configuration, route tables, outbound security group rules, and DNS.
Resolve 502/503 errors and health check failures on ALBs and NLBs — target groups, security groups, and EKS Load Balancer Controller.
Systematically debug a failing CI/CD pipeline — reading step logs, env var validation, authentication errors, and flaky tests.
Fix Docker build failures in CI — registry auth, COPY path errors, build arguments, layer caching, and multi-stage build issues.
Fix Nginx returning 404 when refreshing SPA routes — try_files configuration for React, Vue, and Angular in standalone, Docker, and Kubernetes environments.
Troubleshoot Ingress resources that produce no address, 502 errors, or miss-route traffic — covers controllers, IngressClass, TLS secrets, and backend endpoints.
Fix real client IP not appearing in X-Forwarded-For — covers Nginx proxy headers, AWS ALB/NLB Proxy Protocol, and Kubernetes Ingress Controller ConfigMap settings.
Fix CoreDNS failures, OOMKilled DNS pods, NetworkPolicy blocking port 53, dnsPolicy misconfiguration, and slow external DNS from ndots:5.
Fix pods that can't reach Services inside the cluster — Service selector, port mismatches, NetworkPolicy, kube-proxy health, and cross-namespace FQDN.
Diagnose and stop pod eviction cycles — node memory and disk pressure, missing resource limits, Priority Classes, and eviction prevention.
Node in NotReady state — kubelet crashes, CNI failures, resource pressure, certificate expiry, and container runtime issues.
PVC stuck in Pending — missing StorageClass, no CSI provisioner, access mode mismatches, and zone binding issues.
Service connection failures — empty endpoints, selector mismatches, NetworkPolicy, kube-proxy, and external access issues.
HPA not scaling pods — unknown metrics, missing resource requests, maxReplicas limits, and stabilisation window settings.
Rollout stuck or not progressing — CrashLoopBackOff, ImagePullBackOff, failing readiness probes, and quota limits blocking new pods.
Guides
Task-oriented tutorials that walk through real implementation decisions, not just commands to copy-paste.
Set up IAM Roles for Service Accounts so your pods get fine-grained AWS permissions without static credentials or node-level IAM abuse.
Inject environment-specific values into Kubernetes manifests at deploy time using envsubst — no templating engine required.
Build runbooks that actually get used during incidents — covering scope, ownership, step format, and what to leave out.
Fix GitLab CI timeout errors — the three-level timeout hierarchy, dependency caching, Docker layer caching, and test parallelisation.
Fix hung Jenkins builds — interactive prompts, kubectl without timeouts, input step configuration, and background process handling.
Fix missing or empty environment variables in pipelines — scope issues, Docker env vars, step persistence, and shell quoting problems.
Comparisons
Side-by-side breakdowns of tools and patterns with real trade-offs — so you can make the right call for your specific context.
When peering is enough and when you're paying for simplicity you'll outgrow. Covers scale, cost, and routing complexity.
Read comparison →Full-text search power vs lightweight label-based log aggregation. Choosing between them depends on your query patterns and budget.
Read comparison →Mature flexibility vs tight Git integration. Both have real operational costs — this covers what actually matters in a team context.
Read comparison →About
DevOps Compass exists because most DevOps content falls into one of two camps: too abstract to be useful, or too narrow to give you any real understanding. This site tries to do neither.
Every article is written around a real problem or decision — one you're likely to face on an actual team running actual workloads. The goal isn't to be comprehensive. It's to be genuinely useful when it matters.
No vendor sponsorships that shape recommendations. No padding articles to hit a word count. If something is complex, we say so. If one option is clearly better for most cases, we tell you that too.
More about this site