Asterrr's Handbook

Route 53 routing

Route 53 routing policies, health checks, alias records, nested record trees for high availability, and DNSSEC signing.

Exam tasks: 2.4 (design for reliability: DNS-based failover and traffic distribution), 2.3 (DNSSEC)

The decision: what should decide which endpoint a user gets (a percentage, the user's location, the network latency, the user's IP range, or health alone), and what happens when that endpoint fails?

Choosing a routing policy

The policies

PolicyChooses byHealth checksTypical use
SimpleNothing. Returns every value in the record, and the client picksNoOne endpoint
WeightedRelative weight, 0 to 255 per recordYesBlue/green, canary, A/B testing
LatencyLowest measured latency between the user's network and the record's AWS RegionYesMulti-Region active/active
FailoverPrimary while healthy, otherwise secondaryRequired on primaryActive/passive DR
GeolocationContinent, country or US state of the resolver or EDNS client subnetYesContent licensing, language, data residency
GeoproximityDistance to AWS Regions or coordinates, adjusted by a bias from -99 to 99YesMoving a share of traffic between nearby Regions
Multivalue answerUp to 8 healthy records, chosen at randomYesSimple client-side load spreading
IP-basedCIDR blocks you define in a CIDR collectionYesSending a known ISP or corporate range to a specific endpoint

Geolocation without a default

Geolocation matches only the locations you define. A user from a country without a record gets no answer unless you add a default location record. Latency routing always returns something.

Multivalue is not a load balancer

Multivalue answer returns up to 8 healthy IPs, but it has no connection awareness, and clients cache the answer. When a question asks for even distribution, TLS termination or session stickiness, use an ELB.

Exam signal

Latency and geolocation are easy to confuse. "Best performance for users" is latency. "Users in Germany must be served from Frankfurt" (a compliance or content rule) is geolocation.

Health checks

TypeChecksUse it for
EndpointHTTP, HTTPS or TCP against an IP or domain name. Optionally a string in the first 5,120 bytes of the responsePublic endpoints
CalculatedThe status of up to 256 child health checks, healthy if at least N are healthy"Region is healthy if 2 of its 3 tiers are healthy"
CloudWatch alarmThe state of a CloudWatch alarmPrivate resources, or custom signals such as queue depth or error rate
  • Health checkers run from many locations on the public internet, so they can't reach private IPs. For a private endpoint, publish a metric, alarm on it, and base the health check on the alarm.
  • Standard interval 30 seconds, fast interval 10 seconds. Failure threshold defaults to 3.
  • An endpoint is healthy when more than 18% of checkers report it healthy.
  • Health checks can publish to CloudWatch, so an SNS notification can follow a failover.
8
Maximum healthy records returned by a multivalue answer query
0–255
Range of a weighted record's weight. 0 on every record means equal split
30 s / 10 s
Standard and fast health-check intervals
-99 to 99
Geoproximity bias range
256
Child health checks a calculated health check can include

Alias records vs CNAME

AliasCNAME
At the zone apex (example.com)YesNo. DNS forbids it
TargetsELB, CloudFront, API Gateway, S3 website endpoint, Global Accelerator, Elastic Beanstalk, VPC interface endpoint, another record in the same zoneAny domain name
Query chargesFree for queries to AWS resourcesCharged
TTLTaken from the targetYou set it
Evaluate target healthYes, with no separate health checkNo

Exam signal

"Point the root domain at a load balancer or CloudFront distribution" is always an alias A (or AAAA) record.

Combining policies in record trees

Records can point at other records through alias records, so policies stack. A common global pattern:

  • With evaluate target health turned on at every level, a failure at the bottom propagates upwards. If every latency branch is unhealthy, the failover record serves the maintenance page.
  • Traffic Flow can build the same tree visually as a versioned traffic policy.
  • For DR patterns that use this, see disaster recovery.

DNSSEC signing

DNSSEC lets resolvers verify that answers came from your zone and weren't tampered with.

  • Route 53 signs the zone. You provide the key-signing key (KSK) as a customer managed KMS key that is asymmetric, uses ECC_NIST_P256, and lives in us-east-1. Route 53 manages the zone-signing key itself.
  • To complete the chain of trust, add a DS record at the parent zone (the registrar, or the parent hosted zone).
  • Before signing, lower TTLs and watch for errors with the DNSSECInternalFailure and DNSSECKeySigningKeysNeedingAction CloudWatch metrics.
  • Signing works on public hosted zones. For VPC resolvers, you can turn on DNSSEC validation in Route 53 Resolver.

Disabling DNSSEC in the wrong order

Remove the DS record from the parent first and wait for its TTL to expire, then disable signing. Doing it the other way round makes the zone fail validation and go dark for validating resolvers.

Scenarios

Scenario
A streaming company serves users from us-east-1 and eu-central-1. Users must go to the Region with the best performance. If a Region's application tier fails, its users must move to the other Region, and if both fail, a static status page in S3 must be served. Which configuration meets these requirements?
Scenario
An internal API runs on EC2 instances in private subnets and is reached through a private hosted zone. The team wants Route 53 to fail over to a standby stack in another Region when the API's error rate is high. Health checks against the private IPs always report unhealthy. What should the architect do?

Further reading

On this page