1. Introduction

A misconfigured AWS security group is one of the most common causes of silent connectivity failures. Unlike an application error, a security group deny produces no error message on the client side — traffic simply never arrives. You'll see timeouts, not refusals, which makes the source of the problem easy to misdiagnose as an application bug.

This guide covers how to diagnose, identify, and fix security group misconfigurations across EC2, EKS, RDS, and load balancer scenarios. If you're debugging a node that won't join your EKS cluster or an EKS pod that can't reach the internet, security groups are usually the first place to check.

2. What a Security Group Misconfiguration Means

AWS security groups are stateful virtual firewalls attached to network interfaces. They evaluate inbound and outbound rules to allow or deny traffic. A misconfiguration means one or more rules are missing, use the wrong protocol or port, or reference the wrong source/destination.

The most common symptoms: connection timeouts from EC2 to RDS, pods unable to reach the internet or internal services, EKS nodes failing to reach the control plane, and health checks on load balancers perpetually failing.

3. Common Causes

4. Step-by-Step Diagnosis and Fix

Step 1: Enable VPC Flow Logs to see rejected traffic

The fastest way to confirm a security group block is to look at VPC Flow Logs and filter for REJECT entries. Without Flow Logs, you're guessing.

# Enable flow logs on a VPC (sends to CloudWatch Logs)
aws ec2 create-flow-logs \
  --resource-type VPC \
  --resource-ids vpc-0abc12345 \
  --traffic-type REJECT \
  --log-destination-type cloud-watch-logs \
  --log-group-name /aws/vpc/flowlogs \
  --deliver-logs-permission-arn arn:aws:iam::123456789012:role/FlowLogsRole

# Query CloudWatch Logs Insights for recent REJECTs
# Useful filter: filter @message like /REJECT/ | stats count() by dstPort

Step 2: Inspect security groups on both ends

# List security groups attached to an instance
aws ec2 describe-instances   --instance-ids i-0abc12345   --query 'Reservations[].Instances[].SecurityGroups'

# View all inbound rules for a security group
aws ec2 describe-security-groups   --group-ids sg-0abc12345   --query 'SecurityGroups[].IpPermissions'

# View outbound rules
aws ec2 describe-security-groups   --group-ids sg-0abc12345   --query 'SecurityGroups[].IpPermissionsEgress'

Step 3: Use the AWS Reachability Analyzer

The Reachability Analyzer traces a network path between two resources and tells you exactly which component — security group, NACL, route table, or gateway — is blocking the connection.

# Create a reachability analysis path
aws ec2 create-network-insights-path   --source i-0source12345   --destination i-0dest12345   --protocol TCP   --destination-port 5432

# Start the analysis
aws ec2 start-network-insights-analysis   --network-insights-path-id nip-0abc12345

# Check results (replace with your analysis ID)
aws ec2 describe-network-insights-analyses   --network-insights-analysis-ids nia-0abc12345   --query 'NetworkInsightsAnalyses[].{Reachable:NetworkPathFound,ExplainCode:Explanations[0].ExplanationCode}'

Step 4: Add or fix missing security group rules

# Add inbound rule: allow PostgreSQL from an app security group
aws ec2 authorize-security-group-ingress   --group-id sg-rds-0abc12345   --protocol tcp   --port 5432   --source-group sg-app-0abc12345

# Add inbound rule using a CIDR (less preferred — use SG references where possible)
aws ec2 authorize-security-group-ingress   --group-id sg-0abc12345   --protocol tcp   --port 443   --cidr 10.0.0.0/8

# Add outbound rule (e.g. if you locked down egress)
aws ec2 authorize-security-group-egress   --group-id sg-0abc12345   --protocol tcp   --port 443   --cidr 0.0.0.0/0

Step 5: Check EKS-specific security group requirements

EKS clusters have specific security group requirements for node-to-control-plane communication. If an EKS node is failing to join the cluster, start here:

# Get the cluster security group ID
aws eks describe-cluster   --name my-cluster   --query 'cluster.resourcesVpcConfig.clusterSecurityGroupId'

# Node SG must allow outbound to control plane on 443
# Control plane SG must allow inbound from node SG on 443 and 10250
# Node SG must allow inbound from control plane on 1025-65535 (for webhook traffic)

# Check if the recommended cluster SG rule exists
aws ec2 describe-security-group-rules   --filters Name=group-id,Values=sg-cluster12345   --query 'SecurityGroupRules[?FromPort==`443`]'

Step 6: Check and fix NACLs

# Get the NACL for a subnet
aws ec2 describe-network-acls   --filters Name=association.subnet-id,Values=subnet-0abc12345   --query 'NetworkAcls[].{ID:NetworkAclId,Entries:Entries}'

# NACLs are stateless — you need both inbound and outbound rules.
# Add an inbound ALLOW rule (rule number 100, TCP 5432)
aws ec2 create-network-acl-entry   --network-acl-id acl-0abc12345   --rule-number 100   --protocol tcp   --rule-action allow   --ingress   --cidr-block 10.0.1.0/24   --port-range From=5432,To=5432

5. Verification Steps

# Test connectivity from within the VPC using a bastion or SSM session
# Replace 10.0.2.50:5432 with your target IP and port
nc -zv 10.0.2.50 5432

# Or from a pod in EKS:
kubectl run net-test --image=nicolaka/netshoot --rm -it --restart=Never --   nc -zv my-rds.cluster-abc.us-east-1.rds.amazonaws.com 5432

# Re-run Reachability Analyzer after your fix
aws ec2 start-network-insights-analysis   --network-insights-path-id nip-0abc12345
# Should now return: "NetworkPathFound": true

6. Common Mistakes

7. Prevention Tips

8. FAQ

Why does my connection time out instead of being refused?

A timeout means the traffic was silently dropped — exactly what security groups (and NACLs) do when they deny a connection. A "connection refused" error means the traffic reached the destination but nothing was listening on that port. If you're seeing timeouts, check security groups and NACLs first.

Can I use a security group from one VPC in another VPC?

Not directly. Security group references only work within the same VPC unless you're using VPC Peering with cross-VPC security group referencing enabled, or AWS Resource Access Manager (RAM) to share security groups. In most cross-VPC scenarios, you'll use CIDR ranges or configure separate security groups in each VPC.

I added the inbound rule but it still doesn't work. What else could it be?

Check in order: (1) Is the rule on the correct security group? (2) Is there a NACL on the subnet blocking traffic? (3) Is there a route table entry for the traffic? (4) Is the target service actually running and listening on that port? Use the Reachability Analyzer — it will identify the exact blocking component.

How do security groups interact with Kubernetes network policies?

Security groups operate at the EC2/ENI level (L3/L4), while Kubernetes network policies operate at the pod level using the CNI plugin. On EKS with the VPC CNI, you can optionally use security groups per pod (via Security Groups for Pods feature). Otherwise, security groups apply to the node, and network policies control pod-to-pod traffic within the cluster.

9. Summary

SymptomLikely CauseFix
Connection timeout (no error)Security group or NACL blocking trafficEnable Flow Logs; run Reachability Analyzer
EC2 can't reach RDSMissing inbound rule on RDS SGAdd inbound rule from app SG on port 5432
EKS node won't joinNode SG missing outbound 443 to control planeAdd outbound rule to cluster SG on 443
ALB health checks failingInstance SG not allowing ALB SGAdd inbound from ALB SG on app port
SG rules look right, still blockedNACL on subnet is stateless-blockingAdd NACL allow rules for both directions

Explore More in This Category

Explore more in this category: AWS & Cloud guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.