Route 53 routing
Route 53 routing policies, health checks, alias records, nested record trees for high availability, and DNSSEC signing.
Exam tasks: 2.4 (design for reliability: DNS-based failover and traffic distribution), 2.3 (DNSSEC)
The decision: what should decide which endpoint a user gets (a percentage, the user's location, the network latency, the user's IP range, or health alone), and what happens when that endpoint fails?
Choosing a routing policy
The policies
| Policy | Chooses by | Health checks | Typical use |
|---|---|---|---|
| Simple | Nothing. Returns every value in the record, and the client picks | No | One endpoint |
| Weighted | Relative weight, 0 to 255 per record | Yes | Blue/green, canary, A/B testing |
| Latency | Lowest measured latency between the user's network and the record's AWS Region | Yes | Multi-Region active/active |
| Failover | Primary while healthy, otherwise secondary | Required on primary | Active/passive DR |
| Geolocation | Continent, country or US state of the resolver or EDNS client subnet | Yes | Content licensing, language, data residency |
| Geoproximity | Distance to AWS Regions or coordinates, adjusted by a bias from -99 to 99 | Yes | Moving a share of traffic between nearby Regions |
| Multivalue answer | Up to 8 healthy records, chosen at random | Yes | Simple client-side load spreading |
| IP-based | CIDR blocks you define in a CIDR collection | Yes | Sending a known ISP or corporate range to a specific endpoint |
Geolocation without a default
Geolocation matches only the locations you define. A user from a country without a record gets no answer unless you add a default location record. Latency routing always returns something.
Multivalue is not a load balancer
Multivalue answer returns up to 8 healthy IPs, but it has no connection awareness, and clients cache the answer. When a question asks for even distribution, TLS termination or session stickiness, use an ELB.
Exam signal
Latency and geolocation are easy to confuse. "Best performance for users" is latency. "Users in Germany must be served from Frankfurt" (a compliance or content rule) is geolocation.
Health checks
| Type | Checks | Use it for |
|---|---|---|
| Endpoint | HTTP, HTTPS or TCP against an IP or domain name. Optionally a string in the first 5,120 bytes of the response | Public endpoints |
| Calculated | The status of up to 256 child health checks, healthy if at least N are healthy | "Region is healthy if 2 of its 3 tiers are healthy" |
| CloudWatch alarm | The state of a CloudWatch alarm | Private resources, or custom signals such as queue depth or error rate |
- Health checkers run from many locations on the public internet, so they can't reach private IPs. For a private endpoint, publish a metric, alarm on it, and base the health check on the alarm.
- Standard interval 30 seconds, fast interval 10 seconds. Failure threshold defaults to 3.
- An endpoint is healthy when more than 18% of checkers report it healthy.
- Health checks can publish to CloudWatch, so an SNS notification can follow a failover.
Alias records vs CNAME
| Alias | CNAME | |
|---|---|---|
At the zone apex (example.com) | Yes | No. DNS forbids it |
| Targets | ELB, CloudFront, API Gateway, S3 website endpoint, Global Accelerator, Elastic Beanstalk, VPC interface endpoint, another record in the same zone | Any domain name |
| Query charges | Free for queries to AWS resources | Charged |
| TTL | Taken from the target | You set it |
| Evaluate target health | Yes, with no separate health check | No |
Exam signal
"Point the root domain at a load balancer or CloudFront distribution" is always an alias A (or AAAA) record.
Combining policies in record trees
Records can point at other records through alias records, so policies stack. A common global pattern:
- With evaluate target health turned on at every level, a failure at the bottom propagates upwards. If every latency branch is unhealthy, the failover record serves the maintenance page.
- Traffic Flow can build the same tree visually as a versioned traffic policy.
- For DR patterns that use this, see disaster recovery.
DNSSEC signing
DNSSEC lets resolvers verify that answers came from your zone and weren't tampered with.
- Route 53 signs the zone. You provide the key-signing key (KSK) as a customer managed KMS key that is
asymmetric, uses
ECC_NIST_P256, and lives in us-east-1. Route 53 manages the zone-signing key itself. - To complete the chain of trust, add a DS record at the parent zone (the registrar, or the parent hosted zone).
- Before signing, lower TTLs and watch for errors with the
DNSSECInternalFailureandDNSSECKeySigningKeysNeedingActionCloudWatch metrics. - Signing works on public hosted zones. For VPC resolvers, you can turn on DNSSEC validation in Route 53 Resolver.
Disabling DNSSEC in the wrong order
Remove the DS record from the parent first and wait for its TTL to expire, then disable signing. Doing it the other way round makes the zone fail validation and go dark for validating resolvers.
Scenarios
Latency records give the best performance and, with evaluate target health, drop an unhealthy Region. Nesting them under a failover record adds the S3 page as the last resort. Geolocation follows location, not performance. Multivalue would hand out the S3 endpoint alongside healthy ALBs, and 50/50 weights ignore latency.
Route 53 health checkers are on the internet and can't reach private IP addresses, whatever the security group says. A health check based on a CloudWatch alarm works for private resources and can use a custom signal such as error rate. Multivalue still needs working health checks, and moving the records public exposes an internal name.
Further reading
Private access to S3 content
S3 presigned URLs, CloudFront signed URLs and signed cookies, origin access control, S3 Access Points, Multi-Region Access Points and Block Public Access.
Load balancers
Choosing between Application, Network and Gateway Load Balancers, and how cross-zone balancing, stickiness, TLS, authentication and static IPs work on each.