AWS VPC Fundamentals (Part 1): Routing, Firewalls, and Costs in a Single VPC

When building cloud infrastructure, it is easy to treat a VPC (Virtual Private Cloud) as background configuration: create a few subnets, add a NAT gateway, set some security group rules, and the system works. Then packets stop getting through or the monthly bill spikes. Suddenly, routing, firewalls, and DNS turn out to be intertwined, making it hard to isolate the problem.

The many components inside a VPC may seem disconnected, but they answer three core questions:

  1. Address: What IP address does a resource have, and which address range does it come from? (VPC CIDR, subnets)
  2. Path: Where should a packet go? (Route tables, internet gateways, NAT gateways, VPC endpoints)
  3. Permission: Is the packet allowed through? (Security groups, network ACLs)

Two less visible threads run through all three: names (DNS resolution determines which path a packet takes) and costs (data transfer pricing for each path). This first article in the series breaks down how a single VPC works, providing a clear mental model and a systematic way to troubleshoot it.

 +----------------------------------------------------+
 |                   Single AWS VPC                   |
 |                                                    |
 | +------------------------------------------------+ |
 | | Address: CIDR & Subnets (AZ-a / AZ-b)          | |
 | +------------------------------------------------+ |
 |                         |                          |
 |                         v                          |
 | +------------------------------------------------+ |
 | | DNS: Route 53 Resolver (Base + 2)              | |
 | +------------------------------------------------+ |
 |                         |                          |
 |                         v                          |
 | +------------------------------------------------+ |
 | | Routing: Route Table (LPM) -> IGW / NAT / VPCE | |
 | +------------------------------------------------+ |
 |                         |                          |
 |                         v                          |
 | +------------------------------------------------+ |
 | | Permission: Security Groups & Network ACLs     | |
 | +------------------------------------------------+ |
 |                         |                          |
 |                         v                          |
 | +------------------------------------------------+ |
 | | Billing: Data Transfer & Gateway Costs         | |
 | +------------------------------------------------+ |
 |                                                    |
 +----------------------------------------------------+

1. Address Planning: VPCs, CIDRs, Subnets, and AZs

VPC and Subnet Boundaries

A VPC is a private virtual network scoped to a Region. When creating one, you must specify a primary IPv4 CIDR block, such as 10.0.0.0/16. A subnet is a smaller address range carved out of that VPC CIDR. Each subnet exists in exactly one Availability Zone (AZ).

AZs are groups of data centers with independent power and networking. A failure in one AZ does not affect the others. Because a subnet belongs to a single AZ, its resources share that AZ’s fate.

For high availability across AZs, the network must therefore have a corresponding set of subnets in each target AZ: one each for the public, application, and database tiers. An ALB, Auto Scaling group, and RDS Multi-AZ deployment can then distribute replicas across AZs. Creating subnets alone, without deploying redundant resources, does not provide high availability.

 VPC: 10.0.0.0/16 (Region-level)
 +-- AZ-a
 |   +-- Public Subnet:  10.0.0.0/24
 |   +-- App Subnet:     10.0.32.0/19
 |   +-- DB Subnet:      10.0.128.0/24
 +-- AZ-b
     +-- Public Subnet:  10.0.1.0/24
     +-- App Subnet:     10.0.64.0/19
     +-- DB Subnet:      10.0.129.0/24

AWS requires a subnet prefix length between /28 and /16 (see Subnet CIDR blocks). The number after the slash specifies how many leading bits identify the network; the remaining bits are available for hosts. An IPv4 address has 32 bits, so /24 leaves 8 bits, or 2⁸ = 256 addresses. A /28 leaves just 4 bits, or 2⁴ = 16 addresses.

AWS reserves 5 IP addresses in every subnet: the first 4 and the last 1. In 10.0.0.0/24, these are .0 for the network address, .1 for the VPC router, .2 for DNS, .3 for future use, and .255 for the broadcast address. Subtracting those 5 leaves 256 − 5 = 251 usable addresses on a /24, and only 16 − 5 = 11 on a /28.

All IP-addressed resources in a VPC—EC2, RDS, Lambda network interfaces, interface endpoints, and so on—use ENIs (Elastic Network Interfaces) underneath. IP addresses and security groups attach to ENIs, rather than to the virtual machine itself.

Why Address Planning Needs to Be Right from the Start

Most cloud resource settings can be adjusted later. Subnet addressing is much harder to change: once created, a subnet cannot be expanded dynamically, and running workloads cannot seamlessly switch IP addresses.

Poor initial planning creates three major problems:

A practical starting point is at least /16 for the VPC, /19 or /20 for private subnets hosting dense compute workloads such as containers, and /24 for public subnets that only host load balancers and NAT gateways.

IPv6 and Public IPv4 Costs

Since February 2024, AWS has charged for every public IPv4 address, whether attached to a resource or not. In us-east-1, the rate is $0.005 per hour, or about $3.6 per IP per month; see VPC pricing. The once-common practice of assigning a public IP to every machine now carries a tangible operating cost.

IPv6 offers a way to meet connectivity needs while managing costs. A VPC can run in dual-stack mode and assign globally routable IPv6 addresses.

For outbound-only IPv6 connectivity, use an egress-only internet gateway. It allows connections initiated from inside the VPC and performs no NAT. IPv6 provides enough addresses for each instance to have a globally unique address and communicate externally using that address, without the translation required by IPv4.


2. Routing: Route Tables, IGWs, and NAT Gateways

How Route Tables Choose: Longest Prefix Match

Every subnet must be associated with a route table. That table always contains a system-maintained local route, such as 10.0.0.0/16 → local, that enables communication between all subnets in the VPC.

When a packet’s destination matches multiple routes, the VPC router forwards it using longest prefix match. The prefix is the number after the CIDR slash. A larger number describes a smaller, more specific range, and the router selects the most specific matching route.

Thus, 10.0.0.0/16 → local always takes precedence over 0.0.0.0/0 → nat-xxx. A service-specific prefix list route also takes precedence over the default internet route.

 Packet Destination: 10.0.32.15
 +-- Rule 1: 0.0.0.0/0       -> nat-0123456789abcdef0
     (Prefix /0)
 +-- Rule 2: 10.0.0.0/16     -> local
     (Prefix /16, Matches & Wins ✅)

 Packet Destination: 52.216.0.1 (AWS S3 IP in us-east-1)
 +-- Rule 1: 0.0.0.0/0       -> nat-0123456789abcdef0
     (Prefix /0)
 +-- Rule 2: pl-63a5400a     -> vpce-0123456789abcdef
     (Contains 52.216.0.0/15, Matches & Wins ✅)

Internet Gateway

An internet gateway (IGW) is a fully managed gateway between a VPC and the internet. It offers horizontal scaling and high availability, without bandwidth bottlenecks or a single point of failure.

An IGW serves two roles:

  1. It is the route target for internet traffic in a route table.
  2. It performs one-to-one network address translation (1:1 NAT) between an EC2 instance’s private IP and its corresponding public IPv4 address. The instance’s operating system only ever sees the private IP; the IGW handles all external address translation.

A Public Subnet Is Defined by Its Route

A subnet has no built-in toggle that makes it public. The sole criterion is whether its associated route table has a default route to an internet gateway (0.0.0.0/0 → igw-xxx).

The subnet’s “auto-assign public IPv4” setting only determines whether newly launched instances receive public IPs. An instance with a public IP still cannot connect externally if its route table lacks an IGW entry.

Subnets without an explicit route table association automatically inherit the VPC’s main route table. If someone accidentally adds an IGW route to that table, all subsequent subnets without explicit associations silently become public.

The standard security practice is to keep only the default local route in the main route table and explicitly associate every other subnet with a custom route table.

NAT Gateway: Zonal vs. Regional

Applications in private subnets often need outbound internet access—for downloading packages or calling external APIs—while preventing external clients from connecting directly to them. A NAT gateway provides this through many-to-one SNAT.

Traditional NAT gateways are zonal: they must be deployed in a public subnet in a specific AZ and associated with an Elastic IP. This creates a classic architectural trade-off:

DeploymentAvailabilityCost and Data Transfer Impact
A separate NAT gateway in each AZA failure in one AZ does not affect outbound access in other AZsHourly base charges apply to multiple NAT gateways
One shared NAT gateway for the entire VPCA failure in its AZ immediately cuts internet access for the entire VPCSaves hourly base charges, but cross-AZ traffic incurs charges in both directions ($0.02/GB total for the round trip)

In November 2025, AWS launched Regional NAT Gateway, substantially easing this architectural problem:

Regional NAT has the same hourly rate as Zonal NAT ($0.045 per AZ per hour). Its value lies in removing public subnet configuration risks and simplifying operations. Adding a new AZ requires about 60 minutes of warm-up, and private NAT scenarios are not currently supported.


3. Access Control: Security Groups and Network ACLs

Comparing Their Design and Behavior

Security groups and network ACLs provide defense in depth within a VPC. They differ fundamentally in where they operate and how they track connection state:

AspectSecurity Group (SG)Network ACL (NACL)
Attachment pointENI levelSubnet boundary
Connection stateStateful: responses to allowed inbound traffic are automatically allowedStateless: inbound and outbound packets are checked independently
Rule typesAllow onlyAllow and Deny
EvaluationRules form a union; a match with any rule allows trafficRules are evaluated in ascending rule number; the first match decides
Default behaviorA new SG has no inbound rules and allows all outbound trafficThe default NACL allows everything; a custom NACL initially blocks everything

The Cost of Being Stateless: Ephemeral Ports

Because a NACL is stateless, allowing inbound TCP 443 is only half the job. If an external client connects from a random high port such as 52311, the inbound packet passes. But the server’s response to port 52311 is dropped at the subnet boundary if the outbound rules do not allow the corresponding range.

Custom NACLs therefore need explicit outbound rules for ephemeral ports (commonly 1024–65535).

Referencing SGs to Decouple Access Rules from IP Addresses

In a multi-tier web architecture, the source in an SG rule should reference the source SG ID directly, rather than a fixed CIDR.

 [sg-alb]  Inbound: 443 from 0.0.0.0/0
    |
    v (HTTP 8080)
 [sg-app]  Inbound: 8080 from sg-alb
    |
    v (PostgreSQL 5432)
 [sg-db]   Inbound: 5432 from sg-app

With this approach, any new app instance added by Auto Scaling automatically gains permission to connect to the database as soon as it is assigned sg-app. External traffic also cannot bypass the ALB to reach backend services directly.

When Custom NACLs Make Sense

For most workloads, leave the default NACL open and use SGs as the primary line of defense. Additional NACL configuration is worthwhile in only two situations:

  1. Emergency blocking of a specific source: SGs only allow traffic. To block a malicious IP or CIDR, add a low-numbered Deny rule to the NACL so it takes precedence. However, a NACL has a default limit of just 20 rules per direction, making it unsuitable for a long blocklist. For traffic entering through an ALB or CloudFront, an AWS WAF IP blocklist is usually a better choice.
  2. Defense in depth and separate administrative control: If the database subnet boundary only allows the app subnet’s range, external traffic stays blocked even if someone accidentally opens sg-db to 0.0.0.0/0. Some compliance audits also require this second line of defense to be managed separately from SGs. The trade-off is maintaining rules in both directions, including ephemeral ports.

4. Accessing AWS Services Without Internet Access: VPC Endpoints

When applications in private subnets call managed services such as S3, DynamoDB, SQS, or ECR, they normally reach the services’ public endpoints through a NAT gateway. This traffic does not actually leave the AWS network: according to the VPC FAQ, packets from within AWS to public service endpoints stay on the AWS global backbone.

There are two real costs: a NAT data processing charge for every GB, and the need for an outbound internet path to use these services.

VPC endpoints address both by sending traffic directly to the service without NAT. The two types work very differently. A gateway endpoint adds a route that redirects traffic for a specific service. An interface endpoint places a network interface in a subnet, giving the service a private IP inside the VPC. Its underlying technology is called PrivateLink.

FeatureGateway EndpointInterface Endpoint (PrivateLink)
Supported servicesS3 and DynamoDB onlyMost AWS managed services, S3, and third-party endpoints
ImplementationAdds a prefix list route to a route table (pl-xxx → vpce-xxx)Creates a dedicated ENI with a private IP in a subnet
PricingCompletely free$0.01 per AZ per hour + $0.01 per GB transferred
Security controlsEndpoint policy onlySupports both security groups and an endpoint policy
Hybrid cloud supportNo direct access from on-premises through VPN/DXSupported (it is effectively a private network interface IP inside the VPC)

How an S3 Gateway Endpoint Redirects Traffic

S3 has no private IP inside the VPC. When an application calls S3, DNS still resolves the name to a public IP, such as 52.216.x.x.

Traffic transparently bypasses NAT thanks to longest prefix match (LPM) in the route table. AWS automatically adds a route whose destination is the S3 prefix list. Its ID varies by Region and follows the format pl-xxxxxxxx. The list contains multiple CIDRs used by S3 in that Region, such as 52.216.0.0/15.

These ranges are all more specific than the default route 0.0.0.0/0, so outbound packets are sent directly to the gateway endpoint. No SDK or DNS changes are needed.

How an Interface Endpoint Redirects Traffic: Private DNS

An interface endpoint redirects traffic through private DNS. When enabled, DNS inside the VPC resolves service names such as sqs.us-east-1.amazonaws.com directly to the endpoint’s private IP. Applications and SDKs automatically use PrivateLink without changing any endpoint URLs. Both VPC DNS attributes must be enabled for this to work (see Section 5).

Choosing an Endpoint and Estimating Costs

$0.01×730×3≈$21.9\$0.01 \times 730 \times 3 \approx \$21.9

Enabling PrivateLink for 10 services brings the fixed charge alone to $219 per month. Broad use of interface endpoints is recommended only when data volumes are very large (with NAT charges far above $21.9), or when security requirements demand a network with absolutely no outbound internet access.


5. DNS Inside a VPC: Architecture and Pitfalls

Route 53 Resolver (AmazonProvidedDNS)

Every VPC has a built-in DNS resolver (Route 53 Resolver). Its IPv4 address is always the base address of the primary VPC CIDR plus 2: for 10.0.0.0/16, that is 10.0.0.2. It is also accessible through the global link-local address 169.254.169.253. This is the main reason subnets reserve the .2 address.

Key Attributes and Common Pitfalls

Two VPC attributes control DNS behavior:

AttributeFunctionDefault
enableDnsSupportEnables the built-in Amazon DNS resolution service in the VPCtrue
enableDnsHostnamesAutomatically assigns public DNS names to instances with public IPstrue for the default VPC; false for a custom VPC

IMPORTANT

For a private hosted zone to resolve correctly, or for an interface endpoint to use private DNS, both attributes must be explicitly enabled as true. When creating a custom VPC with Terraform or the AWS CLI, enableDnsHostnames is often overlooked. The result: connections by IP work, but domain names do not resolve.

A custom DHCP option set is another pitfall. If it points instances to your own DNS server, they stop querying AmazonProvidedDNS. Unless that server forwards queries back to 10.0.0.2, private hosted zone and endpoint records will not resolve.

DNS queries to 10.0.0.2 cannot be blocked by SGs or NACLs.

However, traffic to link-local services (Route 53 Resolver DNS, IMDS, NTP, and others) is subject to an aggregate limit of 1024 PPS (packets per second) that cannot be increased; once the quota is reached, the Resolver rejects traffic outright, so highly concurrent workloads with short-lived connections can encounter intermittent DNS query timeouts (i/o timeout). A local DNS cache at the node level, such as NodeLocal DNSCache, is recommended.


6. Network Bills: Data Transfer and Gateway Costs

Here are the main network transfer and gateway rates for us-east-1 (as of September 2026):

ChargeRate
NAT gateway hourly and processing charges$0.045 per AZ per hour + $0.045 per GB processed
Cross-AZ data transfer$0.01 per GB (charged separately in each direction, totaling $0.02 per GB)
Transfer within the same AZFree
Gateway endpointFree
Interface endpoint$0.01 per AZ per hour + $0.01 per GB processed
Public IPv4 usage$0.005 per IP per hour
Internet egressFirst 100 GB per month free; then $0.09 per GB for the first 10 TB

A Practical Cost Comparison: Reading 10 TB from S3

Suppose a batch computing cluster in a private subnet reads 10 TB from S3 each month. The costs vary dramatically by architecture:

One architectural decision makes a difference of nearly $500 per month.


7. Advanced Topic: IP Exhaustion in EKS and the VPC CNI

The Architectural Cost of a VPC-Native CNI

EKS uses the AWS VPC CNI plugin by default. Its design is VPC-native: every Kubernetes pod receives a real private IP directly from the VPC CIDR.

Pods become first-class participants in the VPC. Without overlay encapsulation, they can directly use SGs, Flow Logs, and Direct Connect routes. The trade-off is that growing pod counts rapidly consume the subnet’s IP pool.

In the default secondary IP mode, each EC2 instance type has hardware limits on the number of attached ENIs and the number of IPv4 addresses per ENI. For example, an m5.large supports up to 3 ENIs with 10 IPs each; each ENI reserves its primary IP for the node itself, leaving 3 × 9 = 27 addresses for pods. The max pods formula additionally counts the two host-networked system pods (the VPC CNI and kube-proxy), which do not consume pod IPs, for a ceiling of 29 pods per node.

The warm pool matters even more. To speed up pod startup, the CNI reserves extra IPs for nodes in advance. Those IPs are deducted from the subnet pool before any pod exists.

Symptoms of IP Exhaustion

As the cluster grows and subnet capacity runs out, typical symptoms include:

A /24 subnet, with only 251 usable IPs, can be exhausted within minutes by a few dozen nodes and their warm pool reservations.

Three Mitigations and How to Choose

ApproachMechanismUse Cases and Trade-offs
Prefix delegationAttaches addresses to ENIs in /28 prefixes of 16 IPs eachGreatly increases pod density per node, but requires contiguous subnet blocks; fragmentation can cause InsufficientCidrBlocks
Secondary CIDR + custom networkingAdds a secondary VPC CIDR, such as 100.64.0.0/10, dedicated to podsSeparates pod and node address ranges at the network layer, but requires more configuration and may slightly reduce pods per node
IPv6 EKS clusterCreates a native IPv6 Kubernetes clusterEliminates IPv4 scarcity outright, but requires the entire infrastructure and external dependencies to support IPv6

A practical order of preference:

  1. For a new cluster whose dependencies all support IPv6, an IPv6 cluster is the most durable solution. However, the IP family must be chosen at cluster creation and cannot be changed later.
  2. For clusters that still need IPv4, enable prefix delegation first and use subnet CIDR reservations to preserve contiguous blocks and prevent fragmentation.
  3. Use a secondary CIDR with custom networking only when the primary CIDR is full or corporate routable IP space is limited.

8. Observability and Troubleshooting: Flow Logs and Reachability Analyzer

AWS provides two main tools for troubleshooting VPC networking: one observes actual traffic, the other simulates static paths:


9. Putting It Together: Packet Tracing and Troubleshooting a Three-Tier Architecture

Architecture Specification

Configure a VPC (10.0.0.0/16) across two Availability Zones:

 +--------------------------------------------+
 | VPC: 10.0.0.0/16                           |
 |                                            |
 | [ Public Tier ]                            |
 | AZ-a: 10.0.0.0/24 | AZ-b: 10.0.1.0/24      |
 | Routes: local + 0.0.0.0/0 -> IGW           |
 | SG (sg-alb): Inbound 443 from 0.0.0.0/0    |
 |                                            |
 | [ App Tier ]                               |
 | AZ-a: 10.0.32.0/19 | AZ-b: 10.0.64.0/19    |
 | Routes: local + 0.0.0.0/0 -> Regional NAT  |
 |         + S3 Prefix -> Gateway Endpoint    |
 | SG (sg-app): Inbound 8080 from sg-alb      |
 |                                            |
 | [ DB Tier ]                                |
 | AZ-a: 10.0.128.0/24 | AZ-b: 10.0.129.0/24  |
 | Routes: local only                         |
 | SG (sg-db): Inbound 5432 from sg-app       |
 +--------------------------------------------+

Following Packets Hop by Hop

 User (Internet)
   | (1) DNS lookup -> ALB Public IP
   v
 [ IGW ] -> 1:1 NAT -> Public Subnet (sg-alb allows 443)
   |
   | (2) Local Route
   |     -> App Subnet (sg-app allows 8080 from sg-alb)
   v
 [ App Instance ]
   +-- (3) Local Route
   |     -> RDS Private IP (sg-db allows 5432 from sg-app)
   +-- (4) S3 Prefix List Route
   |     -> S3 Gateway Endpoint (Bypasses NAT, $0)
   +-- (5) Default Route (0.0.0.0/0)
         -> Regional NAT -> IGW -> External API
  1. First hop (client → ALB): The user resolves the ALB’s public IP through DNS. The packet reaches the IGW, which translates the destination to the ALB node’s private IP. It enters the public subnet, where sg-alb allows the request on port 443.
  2. Second hop (ALB → app): The ALB node connects from its private IP to the app instance’s private IP through the VPC’s local route. sg-app verifies that the source is assigned sg-alb and allows port 8080.
  3. Third hop (app → RDS): The app queries the built-in DNS resolver (10.0.0.2) for the RDS hostname and receives the database’s private IP. It follows the local route to the database subnet, where sg-db allows port 5432. If the app and RDS are in different AZs, this hop incurs cross-AZ transfer charges.
  4. Fourth hop (app → S3): The app resolves S3’s public hostname. The route table matches the more specific S3 prefix list route and sends the packet directly to the S3 gateway endpoint, bypassing NAT and incurring no transfer charge.
  5. Fifth hop (app → external API): The destination is a public, internet-routable IP. The packet matches 0.0.0.0/0 and is forwarded to the Regional NAT Gateway, which applies SNAT to the public outbound IP before sending it through the IGW.

Common Symptoms: A Troubleshooting Reference

SymptomPossible Root CauseSuggested Checks
An instance in a public subnet cannot access the internetThe route table is missing 0.0.0.0/0 → IGW, or the subnet is accidentally using the main route tableCheck the subnet’s associated route table; confirm that the instance has a public IP
A private subnet cannot download packages or call external servicesMissing NAT route, or an outbound connectivity problem in the NAT subnet/AZCheck the target of 0.0.0.0/0 in the private route table; use Reachability Analyzer to test the path from the ENI to an external IP
ALB target group health checks keep failingsg-app does not allow the health check port from sg-albCompare the target group’s health check port with the inbound rules in sg-app
All connections time out, or only some clients time outA custom NACL does not allow ephemeral ports, or its range is too narrow (for example, 32768–61000, blocking clients that use other ranges)Look for outbound REJECT records in VPC Flow Logs; add the missing NACL rules for 1024–65535
IP connections work, but private hosted zone names do not resolveThe private hosted zone is not associated with this VPC, or enableDnsHostnames is falseConfirm the VPC association in Route 53; check the attribute with aws ec2 describe-vpc-attribute; use dig to verify that queries go to 10.0.0.2
Traffic still uses NAT after adding an interface endpointPrivate DNS is disabled, so the service hostname still resolves to a public IPRun dig <service-endpoint> on the instance and confirm that the result is a private IP within the VPC
NAT gateway charges spike unexpectedlyHigh-volume access to services such as S3 or ECR lacks a dedicated VPC endpointReview NatGateway-Bytes in Cost Explorer; use Flow Logs to identify high-volume destinations
EKS pods are stuck in ContainerCreatingNo available subnet IPs, or fragmentation prevents prefix delegationCheck the subnet’s remaining available IPs; look for InsufficientCidrBlocks in CNI logs
RDS access has high latency and the bill shows cross-AZ transfer chargesThe app and primary RDS node are in different AZsCheck the AZ distribution of the instances and RDS; review DataTransfer-Regional-Bytes in Cost Explorer

Conclusion and Part 2 Preview

Understanding a single VPC comes down to a focused mental model: CIDRs and subnets define address space, route tables choose paths, security groups and NACLs filter traffic, DNS directs it, and each path determines the cost structure.

For troubleshooting, work through names → paths → permissions in order. Combined with Flow Logs and Reachability Analyzer, this approach resolves most single-VPC networking problems.

A single VPC is only the starting point. As you move into multiple accounts, interconnected VPCs, and hybrid connectivity with on-premises networks, VPC peering, Transit Gateway, Direct Connect, and Route 53 Resolver endpoints enter the picture. Those are the focus of Part 2.