Network Security: VPC, Firewall Rules, and Service Mesh mTLS

Design network security for cloud-native applications using VPCs, network policies, and mutual TLS for service-to-service encryption.

published: reading time: 29 min read author: GeekWorkBench updated: June 17, 2026
Quick Summary

Network security limits which systems can communicate, from cloud VPCs to Kubernetes workloads. This guide compares stateful security groups with stateless NACLs, explains CIDR planning, Kubernetes NetworkPolicy enforcement, and service-mesh mTLS, and covers certificate renewal and zero-trust controls. It also shows how to validate enforcement and prepare for failures such as blocked traffic, expired certificates, and proxy resource pressure.

Introduction

Network security is the discipline of controlling which systems can communicate with which other systems, and under what conditions. It spans from cloud VPC architecture down to individual Kubernetes pods. When done well, it prevents lateral movement during incidents, reduces blast radius from compromised workloads, and keeps sensitive data from leaking to unintended recipients. When done poorly, a single misconfigured security group rule or an overpermissive NACL can cascade into a full production outage or a data breach.

The failures network security prevents are rarely theoretical. A missing security group rule blocks legitimate traffic and causes cascading timeouts. An overly permissive ingress rule exposes internal services to the public internet. A Kubernetes cluster without NetworkPolicy means a single compromised pod can reach every other workload in the cluster. Certificates that are not renewed automatically cause HTTPS outages that affect users directly. Service mesh misconfigurations add latency that surfaces only under production load.

This guide covers the full stack of cloud-native network security: VPC design and CIDR allocation, stateful security groups versus stateless NACLs, Kubernetes NetworkPolicy enforcement, mutual TLS via service mesh (Linkerd and Istio), automated certificate management with cert-manager, and zero-trust architecture principles. By the end, you will be able to design a network security posture that limits blast radius, enforces least-privilege connectivity, and survives common failure scenarios without manual intervention.

When to Use

Istio vs. Linkerd vs. No Service Mesh

Use a service mesh when you have multiple services that need mutual TLS without modifying application code. Linkerd is the better choice when you want simple, low-overhead mTLS with minimal configuration. Istio is better when you need fine-grained traffic control, advanced observability, or multi-cluster federation.

Do not use a service mesh if your services communicate over a dedicated network with no untrusted traffic, or if your team cannot afford the operational overhead. A service mesh adds complexity at every level: debugging, routing, and authentication all become mesh concerns.

The choice between Linkerd and Istio usually comes down to operational tolerance and feature requirements. Linkerd uses its own Rust micro-proxy, while Istio commonly uses Envoy. Both add proxy resources and another layer to inspect during incidents, so benchmark latency and memory use with your own traffic. Linkerd favors a smaller operational surface; Istio offers a broader set of traffic-management and policy controls. Choose based on the features you need and the team’s ability to operate the control plane and data plane.

A common rollout failure is enabling strict mTLS before every client is meshed or otherwise configured to present a trusted identity. Linkerd automatically uses mTLS between meshed pods, while Istio can enforce strict mTLS with PeerAuthentication. Test dependencies and health checks in staging, and verify how each mesh handles traffic from non-meshed clients before enforcing the policy.

cert-manager vs. Cloud-Native Certificate Management

Use cert-manager when you run Kubernetes and want a unified way to manage certificates across cloud and on-premises environments, or when you use Let’s Encrypt or other ACME providers.

Use cloud-native certificate management (AWS ACM, Azure Key Vault, GCP Certificate Manager) when you operate entirely within one cloud and your certificates are primarily for cloud-managed ingresses like ALB or Cloud CDN.

cert-manager is the right choice when your infrastructure spans multiple cloud providers or includes on-premises Kubernetes clusters. It gives you a single control plane for certificate lifecycle management regardless of where your clusters run. The ACME protocol support means you can use Let’s Encrypt at no cost, which is practical for internal services that need publicly trusted certificates. cert-manager also integrates directly with Kubernetes Ingress resources, so certificate provisioning can be automated as part of your existing deployment workflow.

Cloud-native certificate management makes more sense when you are fully committed to one cloud provider and your certificate needs are concentrated at the load balancer layer. AWS ACM handles certificate rotation for ALB and CloudFront, but it does not help you manage certificates inside your cluster for pod-to-pod mTLS. Azure Key Vault with the Azure Key Vault Certificate Injector can bridge this gap, but the setup is more involved than cert-manager. If you are running EKS with an Application Load Balancer and do not need mTLS between services, ACM alone is sufficient and avoids the cert-manager operational surface.

The constraint that trips people up: an ACME HTTP01 challenge requires the validation endpoint to be reachable from the internet. For a service reachable only on a private network but using a domain you control in public DNS, DNS01 can prove control through a public TXT record. An internal-only domain needs a private CA rather than a publicly trusted ACME certificate. DNS01 also requires credentials for a supported DNS provider.

NACLs vs. Security Groups Alone

NACLs add value when you need subnet-wide rules that apply to all resources in a subnet, or when you want explicit deny rules at the network layer (for example, blocking known malicious IP ranges before traffic reaches any security group).

Most workloads do fine with security groups alone. NACLs are worth the additional configuration complexity when you have a specific compliance requirement for network-layer filtering or when multiple security groups need a common deny rule.

NACLs become worth the complexity when two things are true. First, you need to block traffic before it reaches any instance in a subnet. Security groups only filter traffic arriving at an ENI, so a NACL with a deny rule on the subnet edge catches traffic before it consumes ENI-level processing. Second, you need a rule that applies uniformly to all current and future resources in a subnet. Security groups require per-instance attachment, which means new resources do not inherit rules automatically. NACLs cover everything in the subnet CIDR without attachment.

In practice, NACLs are most useful on data subnets where you want a persistent deny boundary that survives security group drift. A common pattern is an ordered NACL rule that allows the app-tier CIDR to reach the database port, followed by a deny rule for other sources. This boundary catches cases where someone accidentally removes the security group restriction on a database instance. Without the NACL layer, that instance becomes publicly accessible the moment its security group rule is deleted or misconfigured.

The operational cost is real. NACLs are stateless, so you must explicitly allow both directions for any conversation. If your app tier at 10.0.1.0/24 connects to your database at 10.0.2.0/24 on port 5432, the data-subnet NACL needs an ingress rule for that connection and an egress rule for the response to the app tier’s ephemeral port. The app-subnet NACL must allow the corresponding return traffic too. Get this wrong and you will spend hours chasing connectivity issues that make no sense. Number your NACL rules (100, 200, 300) so you can trace the evaluation order, and document the expected traffic flows before applying them.

When NOT to Use Network Security Controls

Do not add a control just because it is available. A service mesh can add more operational work than value for a small set of services on a tightly controlled network, especially when the team cannot support proxy upgrades and certificate troubleshooting. NACLs are a poor fit when security groups already express the required boundaries and nobody can maintain the extra stateless return-path rules. NetworkPolicy also depends on a CNI that enforces it; a policy object by itself does not filter traffic.

Keep cloud-managed certificate services for certificates they can actually attach to. They are not a substitute for workload identity or pod-to-pod mTLS. Likewise, avoid blanket network restrictions during early development when service dependencies are still changing; first map the traffic, then apply narrow rules and verify them in staging.

VPC Design and CIDR Allocation

A VPC (Virtual Private Cloud) is your network boundary in the cloud. Design it carefully because changing it later is painful.

Choose CIDR blocks based on expected address use, growth, and connections to other networks. Reserve space for future subnets and availability zones, and avoid ranges that overlap with on-premises networks or other VPCs.

# AWS VPC example
resource "aws_vpc" "main" {
  cidr_block           = "10.0.0.0/16"
  enable_dns_hostnames = true
  enable_dns_support   = true
}

# Subnets across availability zones
resource "aws_subnet" "private" {
  count             = 3
  vpc_id            = aws_vpc.main.id
  cidr_block        = cidrsubnet(aws_vpc.main.cidr_block, 8, count.index)
  availability_zone = data.aws_availability_zones.available.names[count.index]
}

resource "aws_subnet" "public" {
  count                   = 3
  vpc_id                  = aws_vpc.main.id
  cidr_block              = cidrsubnet(aws_vpc.main.cidr_block, 8, count.index + 10)
  availability_zone        = data.aws_availability_zones.available.names[count.index]
  map_public_ip_on_launch = false
}

Segment your VPC into at least three subnet types:

  • Public subnets: Internet-facing load balancers and NAT gateways, with routes through an internet gateway.
  • Private subnets: Application workloads with no direct route from the internet; outbound access may use a NAT gateway or VPC endpoints.
  • Data subnets: Databases and caches with narrowly scoped routes and no general internet egress where practical.

Security Groups and NACLs

Security groups are stateful firewalls attached to instances or ENIs (Elastic Network Interfaces). They are the primary tool for controlling traffic to your workloads.

# Security group for an application tier
resource "aws_security_group" "app" {
  name        = "app-tier"
  description = "Security group for application servers"
  vpc_id      = aws_vpc.main.id

  ingress {
    from_port       = 8080
    to_port         = 8080
    protocol        = "tcp"
    security_groups = [aws_security_group.load_balancer.id]
  }

  # Example dependency: allow only PostgreSQL traffic to the database tier.
  egress {
    from_port       = 5432
    to_port         = 5432
    protocol        = "tcp"
    security_groups = [aws_security_group.database.id]
  }
}

Network ACLs (NACLs) are stateless and operate at the subnet level. Use them as a second layer of defense, for example, blocking all traffic to your data subnets except from specific app subnets.

# NACL for data subnet - only allow traffic from app tier
resource "aws_network_acl" "data" {
  vpc_id = aws_vpc.main.id

  ingress {
    rule_number     = 100
    from_port        = 5432
    to_port          = 5432
    protocol         = "tcp"
    cidr_block       = "10.0.1.0/24"  # App tier subnet
    rule_action      = "allow"
  }

  ingress {
    rule_number  = 200
    cidr_block   = "0.0.0.0/0"
    rule_action  = "deny"
  }

  # Return traffic to the app tier uses its ephemeral source port.
  egress {
    rule_number  = 100
    from_port    = 1024
    to_port      = 65535
    protocol     = "tcp"
    cidr_block   = "10.0.1.0/24"
    rule_action  = "allow"
  }
}

The key difference: security groups are stateful (return traffic is automatically allowed), NACLs are stateless (you must explicitly allow return traffic).

Kubernetes NetworkPolicy Enforcement

In Kubernetes, pods can communicate freely by default. A compromised pod can reach any other pod in the cluster. NetworkPolicy is the Kubernetes-native way to restrict this when the cluster networking implementation enforces the resource.

# Default deny all ingress for a namespace
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-ingress
  namespace: production
spec:
  podSelector: {}
  policyTypes:
    - Ingress

---
# Allow only from frontend to backend
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-frontend-to-backend
  namespace: production
spec:
  podSelector:
    matchLabels:
      app: backend
  policyTypes:
    - Ingress
  ingress:
    - from:
        - podSelector:
            matchLabels:
              app: frontend
      ports:
        - protocol: TCP
          port: 8080

NetworkPolicy resources have no effect unless the cluster networking implementation enforces them. Calico and Cilium provide policy enforcement; Amazon VPC CNI also supports it on supported EKS configurations when the feature is enabled. EKS clusters do not enable this feature by default, so verify the add-on version, configuration, node support, and policy logs before relying on it.

Service Mesh mTLS

When you want encryption and authentication between every service, a service mesh with mutual TLS is the answer. Instead of managing certificates in your application code, the mesh handles it.

flowchart LR
    A[Service A] -->|mTLS| B[Linkerd Proxy]
    B -->|mTLS| C[Linkerd Proxy]
    C --> D[Service B]
    E[Service C] -->|mTLS| F[Linkerd Proxy]
    F -->|mTLS| B
    B --> G[Control Plane<br/>certificates, policies]
    D --> H[Certificate<br/>rotation]

Linkerd automatically enables mTLS between meshed pods. Add the injection annotation to a namespace or workload, then roll out the pods so the proxy is present. For example:

metadata:
  annotations:
    linkerd.io/inject: enabled

Traffic from non-meshed sources may still be accepted as plaintext by default, so use mesh authorization policy when the boundary must reject those connections.

Your services do not change. Linkerd intercepts traffic at the proxy level, verifies certificates, and encrypts communication automatically.

Istio gives you more control but requires more configuration:

# PeerAuthentication policy for strict mTLS
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: default
  namespace: production
spec:
  mtls:
    mode: STRICT

Certificate Management with cert-manager

Whether you are using service mesh mTLS or just securing ingress, certificates need to be managed automatically. cert-manager automates certificate issuance and renewal.

# ClusterIssuer for Let's Encrypt
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-prod
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: ops@example.com
    privateKeySecretRef:
      name: letsencrypt-prod
    solvers:
      - http01:
          ingress:
            class: nginx

Once you have a ClusterIssuer, you can request certificates for any service:

apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: myapp-tls
  namespace: production
spec:
  secretName: myapp-tls
  issuerRef:
    name: letsencrypt-prod
    kind: ClusterIssuer
  dnsNames:
    - myapp.example.com

cert-manager handles renewal automatically. Certificates from Let’s Encrypt expire every 90 days; cert-manager renews them at 30 days by default.

Zero-Trust Network Architecture

Zero-trust means never assuming that a request is safe just because it comes from inside your network. Every request should be authenticated and authorized, regardless of source.

The practical implications:

  1. Identity-based access: Services should have identities (certificates, service accounts) that are verified, not just IP-based access.
  2. Microsegmentation: Each service should only be able to reach the services it needs, nothing more.
  3. Short-lived credentials: Service accounts should use short-lived tokens, not long-lived secrets.

The Secrets Management guide covers how to implement identity-based secrets distribution.

Production Failure Scenarios

Failure Impact Mitigation
Security group rule conflict causing intermittent connectivity Services cannot reach each other, timeouts, partial failures Review effective rules and rule ownership, document expected port ranges, and test after changes
cert-manager failing to renew certificate causing production outage HTTPS becomes unavailable, all TLS connections fail Monitor cert-manager’s cert-ready condition, set up alerts 30 days before expiry, keep a backup certificate
Linkerd mTLS causing latency spikes Service-to-service latency increases, timeouts Profile your services with and without mTLS, use Linkerd’s tap command to identify slow connections, check proxy resource limits
NetworkPolicy misconfigured blocking all traffic to namespace All pods in namespace become unreachable, total outage Apply NetworkPolicy to one pod first, test connectivity before broad rollout, always have a recovery path
NACL overly restrictive blocking legitimate traffic Database or API unreachable from app tier, cascading failures Test NACL changes on a non-production subnet first, use descriptive rule numbers for easy identification

Network Security Observability

Certificate expiration causes complete outages. Set alerts at 60, 30, 7, and 1 day before expiry. If cert-manager reports a certificate as not-ready, investigate immediately — you may have a DNS validation failure or network connectivity issue.

Security group change frequency matters. Teams that modify security groups multiple times per day either have automation problems or unclear ownership. Frequent changes also make auditing harder.

For Linkerd and Istio, watch proxy CPU and memory on each pod. The sidecar adds overhead that catches teams off guard when they have tight resource limits and start getting evicted under load.

Key commands:

# Check cert-manager certificate status
kubectl get certificates -A -o wide

# Monitor certificate expiration
kubectl get certificates -A -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.notAfter}{"\n"}'

# List security group rules in AWS
aws ec2 describe-security-groups --region us-east-1 --query 'SecurityGroups[*].{Name:GroupName,Rules:IpPermissions}'

# Check Linkerd mTLS status
linkerd identity -n production

# Verify NetworkPolicy is applied
kubectl get networkpolicies -A -o wide

Before a rollout, check that:

  • Certificate issuance and renewal succeed, and alerts fire before certificates expire.
  • Security group, NACL, and NetworkPolicy changes are recorded with an owner and a reason.
  • NetworkPolicy enforcement is active in the selected CNI, and denied traffic can be investigated.
  • Service mesh proxies stay within their CPU and memory budgets under expected load.
  • Application teams can distinguish policy denials from routing, DNS, and certificate failures.

Common Anti-Patterns

Relying on the internal network being safe. Internal networks get breached. Compromised containers can reach any other container in the same VPC. Treat internal traffic as untrusted and use mTLS or at minimum application-layer authentication.

Using a certificate that clients do not trust. A self-signed certificate can work when clients deliberately trust it, and private PKI is common for internal services. The operational risk is distributing trust and managing renewal consistently; public CT requirements apply to publicly trusted certificates, not private certificates. Use a managed public CA for public endpoints or a managed internal CA for private workloads.

Not rotating certificates. Certificates expire, get compromised, or need to be replaced after incidents. Automate renewal with cert-manager and test the renewal process before expiry.

Over-trusting the Kubernetes network. Pods can reach any other pod by default. A single compromised workload can become a pivot point to other workloads. A tested default-deny NetworkPolicy can limit lateral movement, but only when the cluster CNI actually enforces the policy.

Attaching overly permissive security group rules as a shortcut. Opening port 0.0.0.0/0 or allowing all traffic from 10.0.0.0/8 “because it’s the internal network” defeats the purpose of security groups. Restrict source CIDRs to the minimum required.

Trade-off Summary

Control Layer Scope Operational Complexity Latency Impact
Security groups Instance-level Low Minimal
NACLs Subnet-level Medium Minimal
VPC endpoints Service-level Medium Path-dependent
PrivateLink / VPC peering Cross-account High Minimal
VPN / Direct Connect On-prem hybrid High Adds encryption overhead
Service mesh (mTLS) Pod-to-pod High Workload-dependent
NetworkPolicy (K8s) Pod-level Medium Minimal

Security and Compliance Notes

Treat the controls in this guide as layers with separate owners. VPC rules restrict network paths, NetworkPolicy narrows pod connectivity, and mTLS verifies workload identity. None of these controls replaces application authorization or encryption of stored data.

For an audit, keep the intended traffic paths, the policies that enforce them, and the review history together. Map each requirement to the control that addresses it, then verify the control in the deployed environment. A passing configuration review is evidence of implementation; it does not by itself establish compliance with PCI DSS, SOC 2, HIPAA, or another framework.

  • Document public entry points and the internal services they can reach.
  • Review broad CIDRs, wildcard selectors, and unrestricted egress on a regular schedule.
  • Retain change records for firewall, certificate, and network policy updates.
  • Test access restrictions and certificate renewal in staging, including failure cases.
  • Assign an owner and response procedure for security alerts and exceptions.

Interview Questions

1. What is the difference between a security group and a NACL in AWS? When would you use both?

Expected answer points:

  • Security groups are stateful — return traffic is automatically allowed without explicit rules
  • NACLs are stateless — you must explicitly allow both inbound and outbound traffic
  • Security groups operate at the instance/ENI level; NACLs operate at the subnet level
  • Use both when you need subnet-wide deny rules (NACLs) plus instance-level source restrictions (security groups)
  • Common pattern: NACLs for known malicious IP blocking, security groups for workload access control
2. How do you design a VPC CIDR allocation plan for a production environment? What considerations affect your choice?

Expected answer points:

  • Size the CIDR for expected address use, growth, availability zones, and connected networks
  • Reserve space for future subnets and avoid overlap with on-premises ranges, peered VPCs, and VPN-connected networks
  • Segment into at least three subnet types: public (load balancers, NAT), private (app workloads), data (databases, caches)
  • Align with availability zones for high availability
  • Avoid overlapping CIDRs across VPC peering, VPN, and on-premises connections
3. Why is the default Kubernetes pod-to-pod network inherently insecure, and how does NetworkPolicy address it?

Expected answer points:

  • By default, all pods in a Kubernetes cluster can reach all other pods — no isolation enforced
  • A compromised pod can pivot laterally to any other workload in the cluster
  • NetworkPolicy is a Kubernetes resource that defines ingress/egress rules per namespace or pod selector
  • Requires policy enforcement from the CNI: Calico and Cilium support it, and EKS Amazon VPC CNI supports it when enabled on supported configurations
  • Best practice: apply a default-deny NetworkPolicy to all production namespaces first, then add allow rules
4. What is mutual TLS (mTLS) and why is it important in a service mesh architecture?

Expected answer points:

  • mTLS is a protocol where both the client and the server authenticate each other with X.509 certificates
  • Unlike regular TLS where only the server presents a certificate, mTLS verifies the client identity as well
  • Prevents unauthorized services from making or receiving connections within the mesh
  • In a service mesh, certificates are managed by the control plane (Linkerd or Istio), not the application code
  • Provides encryption (confidentiality) plus authentication (identity verification) for all service-to-service traffic
5. Linkerd vs. Istio — when would you choose one over the other?

Expected answer points:

  • Linkerd: consider it when automatic mTLS and a smaller operational surface fit your needs; benchmark the proxy with your own workload
  • Istio: choose when you need fine-grained traffic control (mirroring, retries, fault injection), multi-cluster federation, or advanced observability
  • Istio has a steeper learning curve and higher resource consumption
  • Linkerd is CNCF graduated; Istio is CNCF graduated but more complex
  • If you only need mTLS and basic observability, Linkerd may offer a simpler operational fit
6. How does cert-manager work with Let's Encrypt? What happens during certificate renewal?

Expected answer points:

  • cert-manager creates a ClusterIssuer or Issuer resource that references Let's Encrypt as the ACME provider
  • When a Certificate resource is created, cert-manager initiates an ACME challenge (HTTP01 or DNS01) to prove domain ownership
  • Once the challenge is passed, Let's Encrypt issues a certificate stored as a Kubernetes Secret
  • Let's Encrypt certificates expire every 90 days; cert-manager renews them at 30 days by default
  • Renewal is automatic — cert-manager monitors expiry and re-initiates the ACME flow when renewal is due
7. What is a zero-trust network architecture and how does it change how you design network security?

Expected answer points:

  • Zero-trust means no request is trusted simply because it originates inside the network perimeter
  • Every request must be authenticated and authorized regardless of source IP or network location
  • Implications: identity-based access (certificates, service accounts) instead of IP-based trust, microsegmentation, short-lived credentials
  • Internal networks are treated as untrusted — the same rigor applied to public-facing services is applied to east-west traffic
  • Service mesh mTLS, Kubernetes NetworkPolicy, and secrets management are all building blocks of zero-trust
8. What are the risks of using self-signed certificates in a production environment?

Expected answer points:

  • Public clients do not trust a self-signed certificate unless its trust anchor has been installed deliberately
  • Private PKI can be appropriate for internal services, but teams must distribute trust anchors and manage issuance, revocation, and renewal
  • Public Certificate Transparency requirements apply to publicly trusted certificates, not private PKI certificates
  • For public endpoints, use a publicly trusted CA; for internal workloads, use a managed internal CA or an explicitly trusted private CA
9. How do you monitor network security health in a Kubernetes cluster? What metrics and alerts are critical?

Expected answer points:

  • Certificate expiration alerts at 60, 30, 7, and 1 day before expiry — certificate outages are total outages
  • Security group rule change detection — unexpected changes can indicate misconfiguration or compromise
  • Service mesh proxy CPU/memory usage — sidecar overhead catches teams off guard under load
  • Monitor cert-manager's CertificateReady condition — failures indicate DNS validation or ACME challenges failing
  • Linkerd/Istio metrics: request success rates, mTLS handshake latencies, policy violations
10. What is the CNI plugin requirement for NetworkPolicy enforcement on Amazon EKS? Why does this trip people up?

Expected answer points:

  • Amazon VPC CNI supports NetworkPolicy on supported EKS configurations, but enforcement must be enabled and is off by default
  • Check the add-on version, node type, IP family, and current EKS limitations before relying on the feature
  • Calico and Cilium are alternatives when their policy features or platform support better fit the cluster
  • Verify enforcement with policy logs and connectivity tests; a NetworkPolicy object alone does not prove packets are filtered
  • After enabling or installing enforcement, verify policy status and logs, then test both permitted and denied connections
11. Explain the trade-offs between VPC endpoints (private endpoints) and using a NAT gateway for outbound traffic.

Expected answer points:

  • NAT gateway: lets private subnets initiate outbound internet connections; the gateway uses the VPC internet gateway for public destinations
  • VPC endpoints: provide private paths to supported AWS services; the exact path and cost depend on gateway versus interface endpoint type
  • Endpoints can reduce NAT data-processing charges for service traffic, though interface endpoints have their own hourly and data charges
  • PrivateLink is useful for exposing a specific service across VPC or account boundaries without general VPC peering
  • Trade-off: NAT gateway serves general internet egress; endpoints narrow access to supported services and can reduce NAT data processing costs
12. How does a security group rule allowing traffic from 10.0.0.0/8 differ from allowing traffic from a specific security group, and why does it matter?

Expected answer points:

  • CIDR-based rules (10.0.0.0/8) allow any resource within that range — overpermissive, no identity guarantee
  • Security group referencing another security group means only resources attached to that specific security group can connect
  • Security group references scale automatically: new instances attached to the source SG inherit its egress permissions without rule changes
  • Overpermissive CIDR rules are a common anti-pattern — they defeat the purpose of security group least-privilege access
13. What is the difference between ACME HTTP01 and DNS01 challenges in cert-manager? When would you use DNS01?

Expected answer points:

  • HTTP01: cert-manager creates a temporary HTTP resource at `http://domain.com/.well-known/acme-challenge/` — works for publicly reachable domains
  • DNS01: cert-manager creates a DNS TXT record `_acme-challenge.domain.com` — verified by the ACME provider querying DNS
  • Use DNS01 for a privately reachable service when its domain is controlled through public DNS; an internal-only domain needs a private CA
  • DNS01 requires public authoritative TXT records and credentials for a supported DNS provider
  • DNS01 supports wildcard certificates; HTTP01 does not
14. What are the main failure scenarios for cert-manager, and how do you detect them before they cause production outages?

Expected answer points:

  • DNS validation failure: the ACME provider cannot reach the HTTP01 challenge or find the expected public DNS01 TXT record
  • Credential or solver failure: expired DNS-provider credentials, incorrect permissions, or broken challenge routing can prevent issuance or renewal
  • Rate limiting: repeated orders can hit Let's Encrypt limits, including a limit on identical identifier sets; use staging while testing and check current limits before retrying
  • Detection: monitor `kubectl get certificates -A` for Ready=False; check cert-manager logs for Challenge/Certificate events
  • Set up alerts on CertificateReady condition and certificate not-ready events
15. Describe how Linkerd's mTLS works end-to-end, from connection initiation to certificate rotation.

Expected answer points:

  • Each meshed service has a Linkerd micro-proxy that intercepts outbound and inbound TCP traffic
  • On outbound: proxy presents the workload's certificate to the destination's proxy
  • On inbound: proxy verifies the client's certificate before forwarding to the application pod
  • Certificates are issued and rotated by the Linkerd control plane — application code sees encrypted traffic but handles no certificates
  • Control plane uses a trust anchor (root certificate) to issue workload certificates — rotation happens automatically
  • Service accounts are mapped to workload identities in the mesh
16. What are the operational complexities involved in using NACLs alongside security groups? When is the complexity justified?

Expected answer points:

  • NACLs require explicit inbound AND outbound rule definitions (stateless) — twice the rule management surface
  • They operate at the subnet level, so they affect all resources in that subnet simultaneously
  • Complexity is justified in regulated environments (PCI-DSS, HIPAA) where subnet-wide deny rules are required
  • Also justified when you need to block known malicious IP ranges before traffic reaches any security group
  • For most workloads, well-structured security groups alone are sufficient
17. How would you design network security for a three-tier application (web, app, database) deployed in a multi-AZ VPC?

Expected answer points:

  • Public subnets in each AZ: internet-facing load balancers and NAT gateways, routed through an internet gateway as needed
  • Private subnets in each AZ: web and app workloads with security group rules that allow only required tier-to-tier traffic
  • Data subnets in each AZ: database tier accepts port 5432 only from the app-tier security group and has no general internet egress
  • Plan subnets across availability zones and check that all VPC and connected-network CIDRs are non-overlapping
  • NACLs add subnet-level deny rules for known malicious IPs on data subnets
  • Multi-AZ requires careful CIDR allocation per subnet per AZ — plan this in the initial VPC CIDR design
18. What is the blast radius of a compromised pod in a default Kubernetes cluster vs. one with default-deny NetworkPolicy?

Expected answer points:

  • Default (no NetworkPolicy): the compromised pod can reach every other pod in the cluster — it can exfiltrate data from databases, retrieve secrets from other pods, pivot to external services
  • With an enforced default-deny policy, the pod can reach only destinations allowed by the applicable ingress and egress policies
  • Plan required DNS and service traffic first; verify the CNI enforces the policy and test both allowed and denied connections
  • Even with default-deny, a compromised pod can still damage the pod it is allowed to communicate with — defense in depth is necessary
19. How does VPC peering differ from AWS PrivateLink, and when would you choose one over the other?

Expected answer points:

  • VPC peering: creates a direct network connection between two VPCs — traffic uses AWS backbone but appears as from the peer CIDR
  • PrivateLink: exposes a service as a private endpoint within your VPC — the service owner manages it, you do not need a peering connection
  • Choose VPC peering for simple, non-isolated cross-VPC communication within your organization
  • Choose PrivateLink for cross-account service exposure where the service consumer should not have full VPC-level access
  • PrivateLink is more secure for consuming third-party SaaS or AWS services because it does not expose your VPC CIDRs
20. What strategies would you use to reduce the operational overhead of managing security groups at scale across multiple environments and AWS accounts?

Expected answer points:

  • Security group naming conventions and tagging: tag by environment (prod/staging), team, owner for traceability
  • Security group hierarchy: define a "base" security group with common rules (monitoring, logging) and have workload SGs reference it
  • Infrastructure as Code (Terraform, Pulumi): treat security group rules as code — review changes via PR, automate application
  • Avoid rule explosion: use security group references over CIDR rules where possible for automatic scaling
  • Automated drift detection: compare actual SG rules against IaC definitions and alert on drift
  • Cross-account security group sharing via AWS Resource Access Manager (RAM) for shared services

Further Reading

Conclusion

Key Takeaways

  • VPC design sets the foundation: size non-overlapping CIDR ranges for current needs and planned growth
  • Security groups are stateful and instance-level; NACLs are stateless and subnet-level
  • EKS NetworkPolicy enforcement must be explicitly enabled on a supported Amazon VPC CNI configuration, or supplied by another policy-capable CNI
  • Service mesh mTLS offloads certificate management from application code
  • cert-manager automates certificate issuance and renewal across all Kubernetes workloads
  • Zero-trust means authenticating every request regardless of network origin

Network Security Checklist

# 1. VPC uses non-overlapping CIDR ranges with public/private/data subnet segmentation
# 2. Security groups restrict source to specific CIDRs or security groups
# 3. NACLs add subnet-level deny rules for known malicious IPs
# 4. NetworkPolicy enforcement enabled and verified for the selected EKS CNI
# 5. Default-deny NetworkPolicy applied to all production namespaces
# 6. cert-manager ClusterIssuer created for Let's Encrypt
# 7. Certificates auto-renewed 30 days before expiry
# 8. Service mesh mTLS enabled for all production namespaces
# 9. Certificate expiration alerts configured at 60, 30, 7, 1 days

Category

Related Posts

Container Security: Image Scanning and Vulnerability Management

Implement comprehensive container security: from scanning images for vulnerabilities to runtime security monitoring and secrets protection.

#container-security #docker #kubernetes

Deployment Strategies: Rolling, Blue-Green, Canary

Compare and implement deployment strategies—rolling updates, blue-green deployments, and canary releases—to reduce risk and enable safe production releases.

#deployment #devops #kubernetes

Developing Helm Charts: Templates, Values, and Testing

Create production-ready Helm charts with Go templates, custom value schemas, and testing using Helm unittest and ct.

#helm #kubernetes #devops