AWS VPC Fundamentals (Part 1): Routing, Firewalls, and Costs in a Single VPC
When building cloud infrastructure, it is easy to treat a VPC (Virtual Private Cloud) as background configuration: create a few subnets, add a NAT gateway, set some security group rules, and the system works. Then packets stop getting through or the monthly bill spikes. Suddenly, routing, firewalls, and DNS turn out to be intertwined, making it hard to isolate the problem.
The many components inside a VPC may seem disconnected, but they answer three core questions:
- Address: What IP address does a resource have, and which address range does it come from? (VPC CIDR, subnets)
- Path: Where should a packet go? (Route tables, internet gateways, NAT gateways, VPC endpoints)
- Permission: Is the packet allowed through? (Security groups, network ACLs)
Two less visible threads run through all three: names (DNS resolution determines which path a packet takes) and costs (data transfer pricing for each path). This first article in the series breaks down how a single VPC works, providing a clear mental model and a systematic way to troubleshoot it.
+----------------------------------------------------+
| Single AWS VPC |
| |
| +------------------------------------------------+ |
| | Address: CIDR & Subnets (AZ-a / AZ-b) | |
| +------------------------------------------------+ |
| | |
| v |
| +------------------------------------------------+ |
| | DNS: Route 53 Resolver (Base + 2) | |
| +------------------------------------------------+ |
| | |
| v |
| +------------------------------------------------+ |
| | Routing: Route Table (LPM) -> IGW / NAT / VPCE | |
| +------------------------------------------------+ |
| | |
| v |
| +------------------------------------------------+ |
| | Permission: Security Groups & Network ACLs | |
| +------------------------------------------------+ |
| | |
| v |
| +------------------------------------------------+ |
| | Billing: Data Transfer & Gateway Costs | |
| +------------------------------------------------+ |
| |
+----------------------------------------------------+
1. Address Planning: VPCs, CIDRs, Subnets, and AZs
VPC and Subnet Boundaries
A VPC is a private virtual network scoped to a Region. When creating one, you must specify a primary IPv4 CIDR block, such as 10.0.0.0/16. A subnet is a smaller address range carved out of that VPC CIDR. Each subnet exists in exactly one Availability Zone (AZ).
AZs are groups of data centers with independent power and networking. A failure in one AZ does not affect the others. Because a subnet belongs to a single AZ, its resources share that AZ’s fate.
For high availability across AZs, the network must therefore have a corresponding set of subnets in each target AZ: one each for the public, application, and database tiers. An ALB, Auto Scaling group, and RDS Multi-AZ deployment can then distribute replicas across AZs. Creating subnets alone, without deploying redundant resources, does not provide high availability.
VPC: 10.0.0.0/16 (Region-level)
+-- AZ-a
| +-- Public Subnet: 10.0.0.0/24
| +-- App Subnet: 10.0.32.0/19
| +-- DB Subnet: 10.0.128.0/24
+-- AZ-b
+-- Public Subnet: 10.0.1.0/24
+-- App Subnet: 10.0.64.0/19
+-- DB Subnet: 10.0.129.0/24
AWS requires a subnet prefix length between /28 and /16 (see Subnet CIDR blocks). The number after the slash specifies how many leading bits identify the network; the remaining bits are available for hosts. An IPv4 address has 32 bits, so /24 leaves 8 bits, or 2⁸ = 256 addresses. A /28 leaves just 4 bits, or 2⁴ = 16 addresses.
AWS reserves 5 IP addresses in every subnet: the first 4 and the last 1. In 10.0.0.0/24, these are .0 for the network address, .1 for the VPC router, .2 for DNS, .3 for future use, and .255 for the broadcast address. Subtracting those 5 leaves 256 − 5 = 251 usable addresses on a /24, and only 16 − 5 = 11 on a /28.
All IP-addressed resources in a VPC—EC2, RDS, Lambda network interfaces, interface endpoints, and so on—use ENIs (Elastic Network Interfaces) underneath. IP addresses and security groups attach to ENIs, rather than to the virtual machine itself.
Why Address Planning Needs to Be Right from the Start
Most cloud resource settings can be adjusted later. Subnet addressing is much harder to change: once created, a subnet cannot be expanded dynamically, and running workloads cannot seamlessly switch IP addresses.
Poor initial planning creates three major problems:
- Overlapping CIDRs: VPC peering explicitly prohibits connecting two VPCs with overlapping CIDRs. When connecting to corporate data centers, VPNs, or Kubernetes pod networks, overlapping addresses prevent routing from distinguishing whether a destination is local or remote. Packets never reach their intended destination.
- Limited room to grow: A VPC can receive a secondary CIDR later, but this is a workaround. It cannot enlarge an existing, crowded subnet.
- Faster-than-expected IP consumption: A single EC2 instance uses just 1 private IP, but VPC-connected Lambda, ECS with awsvpc networking, EKS pods, interface endpoints, and ALB nodes all consume subnet IPs directly (especially in EKS container environments; see Section 7).
A practical starting point is at least /16 for the VPC, /19 or /20 for private subnets hosting dense compute workloads such as containers, and /24 for public subnets that only host load balancers and NAT gateways.
IPv6 and Public IPv4 Costs
Since February 2024, AWS has charged for every public IPv4 address, whether attached to a resource or not. In us-east-1, the rate is $0.005 per hour, or about $3.6 per IP per month; see VPC pricing. The once-common practice of assigning a public IP to every machine now carries a tangible operating cost.
IPv6 offers a way to meet connectivity needs while managing costs. A VPC can run in dual-stack mode and assign globally routable IPv6 addresses.
For outbound-only IPv6 connectivity, use an egress-only internet gateway. It allows connections initiated from inside the VPC and performs no NAT. IPv6 provides enough addresses for each instance to have a globally unique address and communicate externally using that address, without the translation required by IPv4.
2. Routing: Route Tables, IGWs, and NAT Gateways
How Route Tables Choose: Longest Prefix Match
Every subnet must be associated with a route table. That table always contains a system-maintained local route, such as 10.0.0.0/16 → local, that enables communication between all subnets in the VPC.
When a packet’s destination matches multiple routes, the VPC router forwards it using longest prefix match. The prefix is the number after the CIDR slash. A larger number describes a smaller, more specific range, and the router selects the most specific matching route.
Thus, 10.0.0.0/16 → local always takes precedence over 0.0.0.0/0 → nat-xxx. A service-specific prefix list route also takes precedence over the default internet route.
Packet Destination: 10.0.32.15
+-- Rule 1: 0.0.0.0/0 -> nat-0123456789abcdef0
(Prefix /0)
+-- Rule 2: 10.0.0.0/16 -> local
(Prefix /16, Matches & Wins ✅)
Packet Destination: 52.216.0.1 (AWS S3 IP in us-east-1)
+-- Rule 1: 0.0.0.0/0 -> nat-0123456789abcdef0
(Prefix /0)
+-- Rule 2: pl-63a5400a -> vpce-0123456789abcdef
(Contains 52.216.0.0/15, Matches & Wins ✅)
Internet Gateway
An internet gateway (IGW) is a fully managed gateway between a VPC and the internet. It offers horizontal scaling and high availability, without bandwidth bottlenecks or a single point of failure.
An IGW serves two roles:
- It is the route target for internet traffic in a route table.
- It performs one-to-one network address translation (1:1 NAT) between an EC2 instance’s private IP and its corresponding public IPv4 address. The instance’s operating system only ever sees the private IP; the IGW handles all external address translation.
A Public Subnet Is Defined by Its Route
A subnet has no built-in toggle that makes it public. The sole criterion is whether its associated route table has a default route to an internet gateway (0.0.0.0/0 → igw-xxx).
The subnet’s “auto-assign public IPv4” setting only determines whether newly launched instances receive public IPs. An instance with a public IP still cannot connect externally if its route table lacks an IGW entry.
Subnets without an explicit route table association automatically inherit the VPC’s main route table. If someone accidentally adds an IGW route to that table, all subsequent subnets without explicit associations silently become public.
The standard security practice is to keep only the default local route in the main route table and explicitly associate every other subnet with a custom route table.
NAT Gateway: Zonal vs. Regional
Applications in private subnets often need outbound internet access—for downloading packages or calling external APIs—while preventing external clients from connecting directly to them. A NAT gateway provides this through many-to-one SNAT.
Traditional NAT gateways are zonal: they must be deployed in a public subnet in a specific AZ and associated with an Elastic IP. This creates a classic architectural trade-off:
| Deployment | Availability | Cost and Data Transfer Impact |
|---|---|---|
| A separate NAT gateway in each AZ | A failure in one AZ does not affect outbound access in other AZs | Hourly base charges apply to multiple NAT gateways |
| One shared NAT gateway for the entire VPC | A failure in its AZ immediately cuts internet access for the entire VPC | Saves hourly base charges, but cross-AZ traffic incurs charges in both directions ($0.02/GB total for the round trip) |
In November 2025, AWS launched Regional NAT Gateway, substantially easing this architectural problem:
- A managed resource at the VPC level: AWS manages outbound routing in the VPC infrastructure layer. You do not need to create a public subnet in your VPC to host the NAT gateway; all private subnets point directly to the same Regional NAT ID.
- Higher SNAT connection capacity: A TCP connection is uniquely identified by a four-tuple: source IP, source port, destination IP, and destination port. The finite source port range limits each NAT IP to about 55,000 concurrent connections to the same destination, so high concurrency can exhaust that capacity. Zonal NAT supports up to 8 IPs (2 by default, with a quota increase required), while Regional NAT supports up to 32 per AZ, raising the outbound concurrency limit by a further 4x.
Regional NAT has the same hourly rate as Zonal NAT ($0.045 per AZ per hour). Its value lies in removing public subnet configuration risks and simplifying operations. Adding a new AZ requires about 60 minutes of warm-up, and private NAT scenarios are not currently supported.
3. Access Control: Security Groups and Network ACLs
Comparing Their Design and Behavior
Security groups and network ACLs provide defense in depth within a VPC. They differ fundamentally in where they operate and how they track connection state:
| Aspect | Security Group (SG) | Network ACL (NACL) |
|---|---|---|
| Attachment point | ENI level | Subnet boundary |
| Connection state | Stateful: responses to allowed inbound traffic are automatically allowed | Stateless: inbound and outbound packets are checked independently |
| Rule types | Allow only | Allow and Deny |
| Evaluation | Rules form a union; a match with any rule allows traffic | Rules are evaluated in ascending rule number; the first match decides |
| Default behavior | A new SG has no inbound rules and allows all outbound traffic | The default NACL allows everything; a custom NACL initially blocks everything |
The Cost of Being Stateless: Ephemeral Ports
Because a NACL is stateless, allowing inbound TCP 443 is only half the job. If an external client connects from a random high port such as 52311, the inbound packet passes. But the server’s response to port 52311 is dropped at the subnet boundary if the outbound rules do not allow the corresponding range.
Custom NACLs therefore need explicit outbound rules for ephemeral ports (commonly 1024–65535).
Referencing SGs to Decouple Access Rules from IP Addresses
In a multi-tier web architecture, the source in an SG rule should reference the source SG ID directly, rather than a fixed CIDR.
[sg-alb] Inbound: 443 from 0.0.0.0/0
|
v (HTTP 8080)
[sg-app] Inbound: 8080 from sg-alb
|
v (PostgreSQL 5432)
[sg-db] Inbound: 5432 from sg-app
With this approach, any new app instance added by Auto Scaling automatically gains permission to connect to the database as soon as it is assigned sg-app. External traffic also cannot bypass the ALB to reach backend services directly.
When Custom NACLs Make Sense
For most workloads, leave the default NACL open and use SGs as the primary line of defense. Additional NACL configuration is worthwhile in only two situations:
- Emergency blocking of a specific source: SGs only allow traffic. To block a malicious IP or CIDR, add a low-numbered Deny rule to the NACL so it takes precedence. However, a NACL has a default limit of just 20 rules per direction, making it unsuitable for a long blocklist. For traffic entering through an ALB or CloudFront, an AWS WAF IP blocklist is usually a better choice.
- Defense in depth and separate administrative control: If the database subnet boundary only allows the app subnet’s range, external traffic stays blocked even if someone accidentally opens
sg-dbto0.0.0.0/0. Some compliance audits also require this second line of defense to be managed separately from SGs. The trade-off is maintaining rules in both directions, including ephemeral ports.
4. Accessing AWS Services Without Internet Access: VPC Endpoints
When applications in private subnets call managed services such as S3, DynamoDB, SQS, or ECR, they normally reach the services’ public endpoints through a NAT gateway. This traffic does not actually leave the AWS network: according to the VPC FAQ, packets from within AWS to public service endpoints stay on the AWS global backbone.
There are two real costs: a NAT data processing charge for every GB, and the need for an outbound internet path to use these services.
VPC endpoints address both by sending traffic directly to the service without NAT. The two types work very differently. A gateway endpoint adds a route that redirects traffic for a specific service. An interface endpoint places a network interface in a subnet, giving the service a private IP inside the VPC. Its underlying technology is called PrivateLink.
| Feature | Gateway Endpoint | Interface Endpoint (PrivateLink) |
|---|---|---|
| Supported services | S3 and DynamoDB only | Most AWS managed services, S3, and third-party endpoints |
| Implementation | Adds a prefix list route to a route table (pl-xxx → vpce-xxx) | Creates a dedicated ENI with a private IP in a subnet |
| Pricing | Completely free | $0.01 per AZ per hour + $0.01 per GB transferred |
| Security controls | Endpoint policy only | Supports both security groups and an endpoint policy |
| Hybrid cloud support | No direct access from on-premises through VPN/DX | Supported (it is effectively a private network interface IP inside the VPC) |
How an S3 Gateway Endpoint Redirects Traffic
S3 has no private IP inside the VPC. When an application calls S3, DNS still resolves the name to a public IP, such as 52.216.x.x.
Traffic transparently bypasses NAT thanks to longest prefix match (LPM) in the route table. AWS automatically adds a route whose destination is the S3 prefix list. Its ID varies by Region and follows the format pl-xxxxxxxx. The list contains multiple CIDRs used by S3 in that Region, such as 52.216.0.0/15.
These ranges are all more specific than the default route 0.0.0.0/0, so outbound packets are sent directly to the gateway endpoint. No SDK or DNS changes are needed.
How an Interface Endpoint Redirects Traffic: Private DNS
An interface endpoint redirects traffic through private DNS. When enabled, DNS inside the VPC resolves service names such as sqs.us-east-1.amazonaws.com directly to the endpoint’s private IP. Applications and SDKs automatically use PrivateLink without changing any endpoint URLs. Both VPC DNS attributes must be enabled for this to work (see Section 5).
Choosing an Endpoint and Estimating Costs
- Always enable gateway endpoints for S3 and DynamoDB: They are free and simple to configure, making them a standard baseline.
- Calculate interface endpoint costs at scale: Enabling an interface endpoint for one service in 3 AZs incurs the following monthly fixed charge (see PrivateLink pricing):
Enabling PrivateLink for 10 services brings the fixed charge alone to $219 per month. Broad use of interface endpoints is recommended only when data volumes are very large (with NAT charges far above $21.9), or when security requirements demand a network with absolutely no outbound internet access.
5. DNS Inside a VPC: Architecture and Pitfalls
Route 53 Resolver (AmazonProvidedDNS)
Every VPC has a built-in DNS resolver (Route 53 Resolver). Its IPv4 address is always the base address of the primary VPC CIDR plus 2: for 10.0.0.0/16, that is 10.0.0.2. It is also accessible through the global link-local address 169.254.169.253. This is the main reason subnets reserve the .2 address.
Key Attributes and Common Pitfalls
Two VPC attributes control DNS behavior:
| Attribute | Function | Default |
|---|---|---|
enableDnsSupport | Enables the built-in Amazon DNS resolution service in the VPC | true |
enableDnsHostnames | Automatically assigns public DNS names to instances with public IPs | true for the default VPC; false for a custom VPC |
IMPORTANT
For a private hosted zone to resolve correctly, or for an interface endpoint to use private DNS, both attributes must be explicitly enabled as true. When creating a custom VPC with Terraform or the AWS CLI, enableDnsHostnames is often overlooked. The result: connections by IP work, but domain names do not resolve.
A custom DHCP option set is another pitfall. If it points instances to your own DNS server, they stop querying AmazonProvidedDNS. Unless that server forwards queries back to 10.0.0.2, private hosted zone and endpoint records will not resolve.
DNS queries to 10.0.0.2 cannot be blocked by SGs or NACLs.
However, traffic to link-local services (Route 53 Resolver DNS, IMDS, NTP, and others) is subject to an aggregate limit of 1024 PPS (packets per second) that cannot be increased; once the quota is reached, the Resolver rejects traffic outright, so highly concurrent workloads with short-lived connections can encounter intermittent DNS query timeouts (i/o timeout). A local DNS cache at the node level, such as NodeLocal DNSCache, is recommended.
6. Network Bills: Data Transfer and Gateway Costs
Here are the main network transfer and gateway rates for us-east-1 (as of September 2026):
| Charge | Rate |
|---|---|
| NAT gateway hourly and processing charges | $0.045 per AZ per hour + $0.045 per GB processed |
| Cross-AZ data transfer | $0.01 per GB (charged separately in each direction, totaling $0.02 per GB) |
| Transfer within the same AZ | Free |
| Gateway endpoint | Free |
| Interface endpoint | $0.01 per AZ per hour + $0.01 per GB processed |
| Public IPv4 usage | $0.005 per IP per hour |
| Internet egress | First 100 GB per month free; then $0.09 per GB for the first 10 TB |
A Practical Cost Comparison: Reading 10 TB from S3
Suppose a batch computing cluster in a private subnet reads 10 TB from S3 each month. The costs vary dramatically by architecture:
- Path A (through a NAT gateway):
- Path B (unnecessarily enabling S3 interface endpoints / PrivateLink in 3 AZs):
- Path C (correctly configuring an S3 gateway endpoint):
One architectural decision makes a difference of nearly $500 per month.
7. Advanced Topic: IP Exhaustion in EKS and the VPC CNI
The Architectural Cost of a VPC-Native CNI
EKS uses the AWS VPC CNI plugin by default. Its design is VPC-native: every Kubernetes pod receives a real private IP directly from the VPC CIDR.
Pods become first-class participants in the VPC. Without overlay encapsulation, they can directly use SGs, Flow Logs, and Direct Connect routes. The trade-off is that growing pod counts rapidly consume the subnet’s IP pool.
In the default secondary IP mode, each EC2 instance type has hardware limits on the number of attached ENIs and the number of IPv4 addresses per ENI. For example, an m5.large supports up to 3 ENIs with 10 IPs each; each ENI reserves its primary IP for the node itself, leaving 3 × 9 = 27 addresses for pods. The max pods formula additionally counts the two host-networked system pods (the VPC CNI and kube-proxy), which do not consume pod IPs, for a ceiling of 29 pods per node.
The warm pool matters even more. To speed up pod startup, the CNI reserves extra IPs for nodes in advance. Those IPs are deducted from the subnet pool before any pod exists.
Symptoms of IP Exhaustion
As the cluster grows and subnet capacity runs out, typical symptoms include:
- New pods remain stuck in
ContainerCreatingindefinitely. kubectl describe podreportsFailedCreatePodSandBox: no IP addresses available in network.- Nodes still have ample CPU and memory, but cannot schedule new workloads.
A /24 subnet, with only 251 usable IPs, can be exhausted within minutes by a few dozen nodes and their warm pool reservations.
Three Mitigations and How to Choose
| Approach | Mechanism | Use Cases and Trade-offs |
|---|---|---|
| Prefix delegation | Attaches addresses to ENIs in /28 prefixes of 16 IPs each | Greatly increases pod density per node, but requires contiguous subnet blocks; fragmentation can cause InsufficientCidrBlocks |
| Secondary CIDR + custom networking | Adds a secondary VPC CIDR, such as 100.64.0.0/10, dedicated to pods | Separates pod and node address ranges at the network layer, but requires more configuration and may slightly reduce pods per node |
| IPv6 EKS cluster | Creates a native IPv6 Kubernetes cluster | Eliminates IPv4 scarcity outright, but requires the entire infrastructure and external dependencies to support IPv6 |
A practical order of preference:
- For a new cluster whose dependencies all support IPv6, an IPv6 cluster is the most durable solution. However, the IP family must be chosen at cluster creation and cannot be changed later.
- For clusters that still need IPv4, enable prefix delegation first and use subnet CIDR reservations to preserve contiguous blocks and prevent fragmentation.
- Use a secondary CIDR with custom networking only when the primary CIDR is full or corporate routable IP space is limited.
8. Observability and Troubleshooting: Flow Logs and Reachability Analyzer
AWS provides two main tools for troubleshooting VPC networking: one observes actual traffic, the other simulates static paths:
- VPC Flow Logs (actual traffic observation): Captures IP traffic metadata at the ENI, subnet, or VPC level, including source/destination IPs, ports, protocol, volume, and
ACCEPT/REJECTstatus, but no packet payloads. AREJECTindicates blocking by an SG or NACL. No records at the destination mean the packet never arrived, usually pointing to DNS resolution or routing. - Reachability Analyzer (static path simulation): Analyzes configuration only. Given a source and destination, it simulates a path through route tables, SGs, NACLs, IGWs, and other configurations, determines reachability, and identifies the exact blocking point. It sends no actual network packets and costs $0.10 per analysis.
9. Putting It Together: Packet Tracing and Troubleshooting a Three-Tier Architecture
Architecture Specification
Configure a VPC (10.0.0.0/16) across two Availability Zones:
+--------------------------------------------+
| VPC: 10.0.0.0/16 |
| |
| [ Public Tier ] |
| AZ-a: 10.0.0.0/24 | AZ-b: 10.0.1.0/24 |
| Routes: local + 0.0.0.0/0 -> IGW |
| SG (sg-alb): Inbound 443 from 0.0.0.0/0 |
| |
| [ App Tier ] |
| AZ-a: 10.0.32.0/19 | AZ-b: 10.0.64.0/19 |
| Routes: local + 0.0.0.0/0 -> Regional NAT |
| + S3 Prefix -> Gateway Endpoint |
| SG (sg-app): Inbound 8080 from sg-alb |
| |
| [ DB Tier ] |
| AZ-a: 10.0.128.0/24 | AZ-b: 10.0.129.0/24 |
| Routes: local only |
| SG (sg-db): Inbound 5432 from sg-app |
+--------------------------------------------+
Following Packets Hop by Hop
User (Internet)
| (1) DNS lookup -> ALB Public IP
v
[ IGW ] -> 1:1 NAT -> Public Subnet (sg-alb allows 443)
|
| (2) Local Route
| -> App Subnet (sg-app allows 8080 from sg-alb)
v
[ App Instance ]
+-- (3) Local Route
| -> RDS Private IP (sg-db allows 5432 from sg-app)
+-- (4) S3 Prefix List Route
| -> S3 Gateway Endpoint (Bypasses NAT, $0)
+-- (5) Default Route (0.0.0.0/0)
-> Regional NAT -> IGW -> External API
- First hop (client → ALB): The user resolves the ALB’s public IP through DNS. The packet reaches the IGW, which translates the destination to the ALB node’s private IP. It enters the public subnet, where
sg-alballows the request on port 443. - Second hop (ALB → app): The ALB node connects from its private IP to the app instance’s private IP through the VPC’s local route.
sg-appverifies that the source is assignedsg-alband allows port 8080. - Third hop (app → RDS): The app queries the built-in DNS resolver (
10.0.0.2) for the RDS hostname and receives the database’s private IP. It follows the local route to the database subnet, wheresg-dballows port 5432. If the app and RDS are in different AZs, this hop incurs cross-AZ transfer charges. - Fourth hop (app → S3): The app resolves S3’s public hostname. The route table matches the more specific S3 prefix list route and sends the packet directly to the S3 gateway endpoint, bypassing NAT and incurring no transfer charge.
- Fifth hop (app → external API): The destination is a public, internet-routable IP. The packet matches
0.0.0.0/0and is forwarded to the Regional NAT Gateway, which applies SNAT to the public outbound IP before sending it through the IGW.
Common Symptoms: A Troubleshooting Reference
| Symptom | Possible Root Cause | Suggested Checks |
|---|---|---|
| An instance in a public subnet cannot access the internet | The route table is missing 0.0.0.0/0 → IGW, or the subnet is accidentally using the main route table | Check the subnet’s associated route table; confirm that the instance has a public IP |
| A private subnet cannot download packages or call external services | Missing NAT route, or an outbound connectivity problem in the NAT subnet/AZ | Check the target of 0.0.0.0/0 in the private route table; use Reachability Analyzer to test the path from the ENI to an external IP |
| ALB target group health checks keep failing | sg-app does not allow the health check port from sg-alb | Compare the target group’s health check port with the inbound rules in sg-app |
| All connections time out, or only some clients time out | A custom NACL does not allow ephemeral ports, or its range is too narrow (for example, 32768–61000, blocking clients that use other ranges) | Look for outbound REJECT records in VPC Flow Logs; add the missing NACL rules for 1024–65535 |
| IP connections work, but private hosted zone names do not resolve | The private hosted zone is not associated with this VPC, or enableDnsHostnames is false | Confirm the VPC association in Route 53; check the attribute with aws ec2 describe-vpc-attribute; use dig to verify that queries go to 10.0.0.2 |
| Traffic still uses NAT after adding an interface endpoint | Private DNS is disabled, so the service hostname still resolves to a public IP | Run dig <service-endpoint> on the instance and confirm that the result is a private IP within the VPC |
| NAT gateway charges spike unexpectedly | High-volume access to services such as S3 or ECR lacks a dedicated VPC endpoint | Review NatGateway-Bytes in Cost Explorer; use Flow Logs to identify high-volume destinations |
| EKS pods are stuck in ContainerCreating | No available subnet IPs, or fragmentation prevents prefix delegation | Check the subnet’s remaining available IPs; look for InsufficientCidrBlocks in CNI logs |
| RDS access has high latency and the bill shows cross-AZ transfer charges | The app and primary RDS node are in different AZs | Check the AZ distribution of the instances and RDS; review DataTransfer-Regional-Bytes in Cost Explorer |
Conclusion and Part 2 Preview
Understanding a single VPC comes down to a focused mental model: CIDRs and subnets define address space, route tables choose paths, security groups and NACLs filter traffic, DNS directs it, and each path determines the cost structure.
For troubleshooting, work through names → paths → permissions in order. Combined with Flow Logs and Reachability Analyzer, this approach resolves most single-VPC networking problems.
A single VPC is only the starting point. As you move into multiple accounts, interconnected VPCs, and hybrid connectivity with on-premises networks, VPC peering, Transit Gateway, Direct Connect, and Route 53 Resolver endpoints enter the picture. Those are the focus of Part 2.