DNS debugging · 24 min read
When DNS Fails: A Debugging Guide for Authoritative Nameserver Resolution
Learn how to trace a broken lookup from resolver to authoritative nameserver, isolate delegation, DNSSEC, and TTL failures, and fix the root cause instead of guessing.
Effective troubleshooting DNS resolution errors requires isolating whether a failure originates at the local stub resolver, the recursive caching layer, the parent delegation registry, or the authoritative nameservers themselves. By bypassing intermediate caches and querying each layer directly using tools like dig and delv, engineers can differentiate an expired DNSSEC signature from an invalid glue record or a stale negative cache entry in seconds.
When an application fails to connect to an external API or users report intermittent service outages, DNS is frequently the primary suspect. Yet engineering teams often waste critical incident minutes flushing local caches or blindly toggling zone records without diagnosing the root cause. This guide outlines an authoritative, step-by-step workflow for troubleshooting DNS resolution errors across modern cloud architectures, providing the exact commands, header flags, and decision paths needed to restore resolution quickly.
The 60-Second Triage: Fix the Most Common DNS Resolution Errors First
When an incident begins, your first priority is narrowing the failure domain. Avoid changing infrastructure or registrar configurations during initial triage. Instead, establish whether the failure is global or local, whether it affects recursive resolvers or authoritative nameservers, and which return code (RCODE) the resolver returns.
Start by extracting the exact fully qualified domain name (FQDN) and record type requested by your application. Run two diagnostic queries: one to a major public recursive resolver (such as Cloudflare at 1.1.1.1 or Google at 8.8.8.8) and one directly to your assigned authoritative nameserver.
# Query recursive resolver
dig @1.1.1.1 api.example.com A +noall +comments +answer
# Query authoritative nameserver directly
dig @ns1.dnscove.com api.example.com A +norecurse +noall +comments +answer
Based on the return codes and network behavior of these commands, classify the outage into one of four primary failure buckets:
- NXDOMAIN (Non-Existent Domain): The authoritative server responded affirmatively that the record does not exist in the zone. Direct fix: Verify typos in the record name, inspect zone apex formatting, or confirm whether an automated CI/CD pipeline inadvertently deleted the record.
- SERVFAIL (Server Failure): The recursive resolver encountered a failure during resolution, frequently due to a DNSSEC validation breakdown, lame delegation, or authoritative timeout. Direct fix: Query the resolver using the
+cd(Checking Disabled) flag. If the query succeeds with+cd, the failure is a broken DNSSEC trust chain. If it still fails, inspect parent delegation and upstream nameserver reachability. - Timeout / Connection Refused: The client received no packet back on UDP/TCP port 53. Direct fix: Verify network routing, firewall state tables, and security group egress/ingress rules. DNSCove runs two unicast authoritative nameservers (ns1 in NYC, ns2 in Frankfurt), not an anycast network. A timeout on one nameserver should be tested against the other before assuming a global outage.
- Wrong Answer / Stale Record: The returned IP or target differs from the expected value. Direct fix: Inspect the record TTL and the zone's authoritative source. If the authoritative server returns the correct IP but recursive resolvers return old data, the resolver is serving valid cached data until its TTL expires.
When evaluating responses, pay close attention to the response header flags. The aa (Authoritative Answer) flag confirms the responding server holds the zone data rather than serving it from cache. The ad (Authenticated Data) flag indicates the recursive resolver successfully validated the DNSSEC cryptographic chain. Reading these flags upfront prevents engineers from modifying authoritative zone files when the underlying issue is an upstream recursive resolver misconfiguration.
How a Lookup Actually Travels: Resolver, Root, TLD, and Authoritative Layers
To master troubleshooting DNS resolution errors, you must understand the path a DNS packet traverses. When an application initiates a network socket, the operating system's stub resolver forwards the request to an upstream recursive resolver. From there, the recursive resolver performs an iterative discovery process across the DNS hierarchy, as defined by standards organizations like IANA.
- The Root Layer: The recursive resolver consults its built-in root hints file to query one of the 13 root nameserver clusters (
a.root-servers.netthroughm.root-servers.net) for the Top-Level Domain (TLD) authoritative servers. - The TLD Layer: The root nameserver returns a referral containing the NS records for the target TLD (such as
.comor.org). The resolver queries the TLD server for the domain's assigned authoritative nameservers. - The Authoritative Layer: The TLD server returns a referral containing the NS records (and necessary glue records) configured at your domain registrar. Finally, the resolver queries these authoritative nameservers directly for the requested child record.
A critical architectural distinction is the difference between recursion and iteration. Stub resolvers request recursion from recursive resolvers by setting the RD flag in the query header. However, as outlined in the core protocol architecture of RFC 1034 and RFC 1035, authoritative nameservers are designed to answer iteratively, returning the authoritative answer, an explicit NXDOMAIN, or a referral to another nameserver set. If an authoritative nameserver attempts to resolve upstream records recursively, it risks severe operational vulnerabilities, including exposure to cache poisoning and amplification vectors.
You can trace this entire hierarchy in real time using dig +trace:
dig +trace api.example.com A
When reading dig +trace output, examine each transition section:
; <<>> DiG 9.18.18-1 <<>> +trace api.example.com A
;; global options: +cmd
. 518400 IN NS a.root-servers.net.
...
;; Received 525 bytes from 127.0.0.53#53(127.0.0.53) in 12 ms
com. 172800 IN NS a.gtld-servers.net.
...
;; Received 1173 bytes from 198.41.0.4#53(a.root-servers.net) in 24 ms
example.com. 172800 IN NS ns1.dnscove.com.
example.com. 172800 IN NS ns2.dnscove.org.
;; Received 654 bytes from 192.5.6.30#53(a.gtld-servers.net) in 32 ms
api.example.com. 300 IN A 192.0.2.10
example.com. 3600 IN NS ns1.dnscove.com.
example.com. 3600 IN NS ns2.dnscove.org.
;; Received 120 bytes from 198.51.100.1#53(ns1.dnscove.com) in 18 ms
If the trace stops at the TLD level without printing the final authoritative response, the break exists between the registrar's delegation and the authoritative servers. One frequent culprit is missing or invalid glue records. When a domain's nameservers sit under the domain itself (for example, example.com using ns1.example.com), the parent TLD zone must provide both the NS record and an A/AAAA record (the "glue") pointing to that nameserver's IP address. If the glue record is missing or contains an obsolete IP, resolvers cannot reach the nameserver, creating a circular resolution lock that causes a total lookup failure.
Another overlooked layer in troubleshooting DNS resolution errors is negative caching. Defined in RFC 2308, negative caching dictates that when a domain or record returns NXDOMAIN or NODATA (a domain exists, but the requested record type does not), recursive resolvers cache that negative answer for the duration specified in the zone's Start of Authority (SOA) minimum TTL field. If you create a missing record in your authoritative dashboard during an outage, external users may continue receiving NXDOMAIN until that negative cache window expires.
DNS Lookup Failure Analysis: Reading dig, drill, and Resolver Output Like a Debugger
Systematic DNS lookup failure analysis depends on parsing DNS response packets with the precision of a protocol debugger. Tools like dig (Domain Information Grok) expose the raw DNS wire format, divided into clear semantic sections.
Consider the anatomy of a comprehensive dig response:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 48214
;; flags: qr aa rd ra; QUERY: 1, ANSWER: 1, AUTHORITY: 2, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags:; udp: 1232
;; QUESTION SECTION:
;api.example.com. IN A
;; ANSWER SECTION:
api.example.com. 300 IN A 192.0.2.10
;; AUTHORITY SECTION:
example.com. 3600 IN NS ns1.dnscove.com.
example.com. 3600 IN NS ns2.dnscove.org.
;; Query time: 14 msec
;; SERVER: 198.51.100.1#53(ns1.dnscove.com)
;; WHEN: Mon Oct 05 14:22:10 UTC 2026
;; MSG SIZE rcvd: 124
When conducting DNS lookup failure analysis, parse the response systematically:
- Header Status (
RCODE): The outcome code returned by the server.NOERRORindicates successful processing.NXDOMAINmeans the name does not exist.SERVFAILsignifies an internal resolver error or DNSSEC validation fault.REFUSEDindicates the nameserver refused to answer (often due to access control or query restrictions).FORMERRmeans the nameserver could not understand the query format. - Header Flags:
qr(Query/Response): Set if the packet is a response.aa(Authoritative Answer): Confirms the answering nameserver is the authoritative source for the zone.rd(Recursion Desired): Set by the client if recursion was requested.ra(Recursion Available): Set by the server if it supports recursive queries.ad(Authentic Data): Set by validating resolvers when DNSSEC signatures are verified.cd(Checking Disabled): Used by clients to instruct resolvers not to perform DNSSEC validation.tc(Truncated Response): Indicates the response exceeded the maximum UDP packet size and was truncated, signaling the client must retry over TCP.
- QUESTION SECTION: Echoes back the requested name, class (usually
INfor Internet), and record type. - ANSWER SECTION: Contains the resource records matching the question.
- AUTHORITY SECTION: Lists the authoritative nameservers for the domain or the SOA record in the event of negative responses.
- ADDITIONAL SECTION / OPT PSEUDOSECTION: Houses Extension Mechanisms for DNS (EDNS0) parameters, such as advertised buffer sizes (e.g.,
udp: 1232) and TSIG or OPT records.
To accelerate your analysis during an incident, employ these specialized command patterns:
# Query an authoritative server without recursion to test authoritative state
dig +norecurse @ns1.dnscove.com api.example.com A
# Query with DNSSEC records requested (+dnssec)
dig +dnssec @1.1.1.1 api.example.com A
# Query disabling DNSSEC validation (+cd) to test for DNSSEC breakages
dig +cd @1.1.1.1 api.example.com A
# Machine-readable output for shell scripts and health checks
dig +short @1.1.1.1 api.example.com A
Note that querying ANY records (e.g., dig ANY example.com) is no longer a reliable diagnostic technique. In accordance with RFC 8482, many modern authoritative nameservers either refuse ANY queries outright or return a synthetic HINFO response to prevent nameservers from being abused in reflection-amplification DDoS attacks. When running diagnostic scripts, query specific record types (A, AAAA, CNAME, TXT, MX) rather than relying on ANY.
The following decision matrix maps observed output directly to probable causes and subsequent diagnostic commands:
| Observed Header / RCODE | Resolver vs Authoritative | Probable Root Cause | Next Diagnostic Command |
|---|---|---|---|
SERVFAIL |
Resolver | DNSSEC validation failure, expired RRSIG, or lame delegation. | dig +cd @resolver domain.com A |
REFUSED |
Authoritative | Server does not host zone; query sent to unconfigured nameserver. | dig @tld-server domain.com NS |
NXDOMAIN |
Authoritative | Record missing, typo in FQDN, or invalid zone apex entry. | dig @auth-ns domain.com SOA |
tc flag set (Truncation) |
Both | Response exceeded UDP buffer (typically 1232 bytes); TCP port 53 blocked. | dig +tcp @auth-ns domain.com A |
| Timeout (no response) | Authoritative | Firewall blocking UDP port 53; network blackhole; routing failure. | nc -zvu auth-ip 53 && nc -zv auth-ip 53 |
Authoritative Nameserver Debugging: Delegation, Glue, and NS Record Mismatches
When an outage is isolated to the authoritative tier, authoritative nameserver debugging focuses on delegation consistency between parent and child zones. A domain's delegation begins at the parent TLD registry. If the TLD nameservers point to one set of authoritative nameservers while the child zone's own NS resource records advertise another, resolvers will encounter intermittent resolution failures, split-brain states, or latency spikes.
To inspect the parent zone's delegation directly, query the TLD nameservers rather than a public caching resolver:
# Query the TLD servers directly for your domain's delegation
dig @a.gtld-servers.net example.com NS +noall +authority +additional
Examine the returned AUTHORITY SECTION. These are the NS records published at the registrar. Next, query each delegated authoritative nameserver directly for its own NS record set:
dig @ns1.dnscove.com example.com NS +noall +answer
dig @ns2.dnscove.org example.com NS +noall +answer
If the TLD nameservers return nameserver hosts that do not match the answer returned by the authoritative nameservers, your domain is suffering from delegation skew. Customer zones are delegated to the shared ns1.dnscove.com / ns2.dnscove.org nameservers; per-customer vanity or white-label nameservers are not supported in v1. Delegation checks should target those shared names. When auditing your infrastructure against configuration errors, consult our DNS zone file audit checklist to ensure parity between your registrar delegation and your authoritative records.
A severe consequence of delegation skew is "lame delegation." A delegation is considered lame when the parent zone delegates authority to a nameserver that is not actually configured to answer authoritatively for that domain. When a recursive resolver queries a lame nameserver, the server typically returns REFUSED or an empty response without the aa flag. If all delegated nameservers are lame, all lookups fail with SERVFAIL across all validating resolvers worldwide.
You can identify lame delegation across your nameserver fleet using this quick inspection script:
for ns in ns1.dnscove.com ns2.dnscove.org; do
echo "--- Testing $ns ---"
dig +norecurse @$ns example.com SOA +noall +comments +answer
done
Check the comments line of each response. If the flags list lacks aa, or if the status is REFUSED, that specific nameserver is operating lamely. Furthermore, compare the answers returned by each server in your authoritative pool. In a dual-nameserver setup, ns1 and ns2 must return identical record sets and serial numbers in their SOA records. If one nameserver returns an updated record while the other returns stale data, client lookups will alternate between success and failure depending on which nameserver the recursive resolver selects.
DNSSEC Validation Failures: When SERVFAIL Means a Broken Chain of Trust
Under modern DNS operational standards, DNSSEC validation failures account for many unexplained SERVFAIL errors. When DNSSEC is enabled, recursive resolvers validate cryptographic signatures on every response before returning data to the client, as specified in RFC 4033. If any link in the cryptographic chain fails, the resolver rejects the response entirely and returns SERVFAIL to protect the client from potential spoofing or cache poisoning.
The DNSSEC chain of trust operates hierarchically:
- The Root zone signs the TLD's public Key Signing Key (KSK) hash, stored as a Delegation Signer (
DS) record in the Root zone. - The TLD zone signs the child zone's public KSK hash, stored as a
DSrecord in the TLD zone. - The child zone's KSK signs the child zone's Zone Signing Key (ZSK), stored as a
DNSKEYrecord at the child apex. - The child zone's ZSK generates Resource Record Signatures (
RRSIG) for every individual record set (RRset) in the child zone.
If you suspect a DNSSEC validation failure, confirm it immediately using dig with the Checking Disabled (+cd) flag:
# Standard query returning SERVFAIL
dig @1.1.1.1 api.example.com A
# Query bypassing DNSSEC validation
dig @1.1.1.1 api.example.com A +cd
If the standard query returns SERVFAIL while the query with +cd returns NOERROR and the correct A record in the answer section, your issue is unequivocally a DNSSEC validation failure. To pinpoint the exact failure point in the cryptographic chain, use delv (Domain Entry Lookup and Validation):
delv @1.1.1.1 api.example.com A
The output of delv will report the precise cryptographic fault, such as validation failure <api.example.com. A IN>: expired signature or no valid RRSIG found. Common root causes of DNSSEC validation failures include:
- Orphaned DS Records: A domain was migrated to a new DNS provider, but the old provider's DS record remains published at the registrar. Resolvers attempt to validate the new provider's signatures against the old provider's keys, resulting in near-universal SERVFAIL errors on all validating resolvers.
- Expired RRSIG Signatures: Authoritative nameservers must regularly re-sign RRsets before their cryptographic validity window lapses. If a zone signing system stalls, signatures expire, and validating resolvers immediately reject the zone.
- Algorithm Mismatches: The
DSrecord published at the registrar specifies a digest algorithm or key algorithm that does not match the actualDNSKEYactive in the zone. - Validator Clock Skew: DNSSEC signatures include strict UTC start and expiration timestamps (
inceptionandexpirationin the RRSIG). If a client or recursive resolver's internal NTP clock drifts, valid signatures will be judged as not yet valid or already expired.
Modern cryptographic configurations also impact authenticated denial of existence. Under RFC 9276, security standards recommend configuring NSEC3 records with 0 additional hash iterations and an empty salt to reduce validator CPU exhaustion and eliminate vulnerability to NSEC3 iteration attacks, as detailed in RFC 9276. Furthermore, registrars and registries increasingly support RFC 7344 and RFC 8078 for automated trust maintenance via CDS (Child DS) and CDNSKEY records, as specified in RFC 7344.
DNSCove signs zones with DNSSEC. It is per zone, enabled with one click, and included on every plan including Free at no extra charge. Algorithm 13 (ECDSA P-256/SHA-256), NSEC3 with RFC 9276 parameters (0 iterations, no salt), and CDS/CDNSKEY published per RFC 7344/8078 for registrar automation. Zone signing keys are held in the control plane under AWS KMS and are never present on the authoritative nameservers. Signatures are refreshed automatically before expiry. Zone-signing keys roll automatically on a 90-day pre-publish schedule, which requires nothing from the customer. The key-signing key is rolled on operator demand rather than on a schedule, because a KSK roll requires a DS change at the registrar. If you need step-by-step instructions on setting up automated trust delegation, follow our DNSSEC management guide.
TTL, Caching, and Propagation: Why Your Fix Has Not Taken Effect Yet
A frequent frustration during DNS incident response is the observation that a corrected record "has not propagated." In standard DNS architecture, authoritative nameservers do not "push" updates to global resolvers. Instead, recursive resolvers pull records on demand and cache them based on the record's Time-to-Live (TTL) header.
Resolution latency during an update is governed by a multi-tiered caching topology:
- Application and Browser Cache: Chromium-based browsers, Node.js runtime pools, and Java Virtual Machines (JVMs) often maintain their own in-process DNS caches, sometimes ignoring TTL values entirely unless explicitly configured.
- Operating System Stub Cache: Daemons such as
systemd-resolvedon Linux ormDNSResponderon macOS store answers locally to minimize outbound network requests. - Local Network / Corporate Forwarders: Enterprise firewalls and local branch office routers cache records across corporate devices.
- Public / ISP Recursive Resolvers: Upstream resolvers (e.g., your ISP, Cloudflare, Google Public DNS) enforce record TTLs independently across their geographically distributed points of presence.
You can observe TTL decay in real time by executing repeated queries against a recursive resolver:
dig @1.1.1.1 api.example.com A +noall +answer
# Output: api.example.com. 287 IN A 192.0.2.10
# Run again 30 seconds later:
dig @1.1.1.1 api.example.com A +noall +answer
# Output: api.example.com. 257 IN A 192.0.2.10
The TTL value decrements with each query until reaching zero, at which point the resolver discards the cached entry and fetches the authoritative record anew. If you need to verify whether an authoritative change has taken effect without waiting for local caches to clear, execute a cache-busting query against a dynamic label (if wildcards are active) or query the authoritative nameservers directly.
Negative caching TTL introduces additional propagation delays. When a user queries a non-existent subdomain (such as during a failed service rollout), the resolver caches the resulting NXDOMAIN based on the minimum TTL defined in the zone's SOA record:
example.com. 3600 IN SOA ns1.dnscove.com. hostmaster.example.com. (
2026100501 ; Serial
7200 ; Refresh
3600 ; Retry
1209600 ; Expire
300 ; Negative Cache TTL / Minimum
)
In this example, the negative caching TTL is 300 seconds. Even if you create the missing record immediately, any resolver that queried the missing record during the failure will refuse to query the authoritative server again until those 300 seconds elapse.
Architects frequently ask what TTL value to assign to production records. Lower TTLs (such as 60 or 300 seconds) provide failover agility, enabling rapid record repointing during an active incident. Higher TTLs (such as 3600 or 86400 seconds) reduce overall query volume, increase client lookup speed via higher cache hit rates, and protect downstream availability if an authoritative server experiences temporary disruption. DNSCove uses fixed-cost pricing rather than per-zone or per-query metering. Consequently, TTL tuning is not a cost decision. You can evaluate our subscription plans and feature inclusions directly on the pricing page.
Edge Cases: Apex Records, CNAME Conflicts, and Split-Horizon Views
Certain infrastructure configurations introduce DNS edge cases that behave unpredictably under standard debugging procedures. The most prevalent of these involves zone apex aliasing.
According to RFC 1034 (Section 3.6.2), if a CNAME record exists at a node, no other data records of any type may coexist at that same node. Because a zone apex (e.g., example.com) is required by protocol to contain SOA and NS records, standard CNAME records cannot legally be placed at the apex. Attempting to force a CNAME at the zone apex violates the DNS specification and causes BIND, Unbound, and recursive resolvers worldwide to drop or misroute queries for the entire zone. Major cloud providers established proprietary mechanisms like AWS Route 53 Alias to address this constraint, as documented in the AWS Route 53 Developer Guide.
DNSCove supports apex ALIAS records (CNAME-at-apex flattening, like Route53 Alias) with serve-stale protection. Under ALIAS flattening, the authoritative nameserver queries the upstream target CNAME internally and synthesizes standard A and AAAA records directly in the authoritative response, keeping the zone fully compliant with RFC 1034 while allowing apex domains to point to dynamic cloud load balancers. For a comprehensive architectural breakdown of apex routing patterns, read our deep-dive on CNAME flattening for apex domains.
A related edge case is the CNAME coexistence collision at subdomains. As clarified in RFC 2181 (Section 10.1), administrators must not assign any other record type to a node that contains a CNAME, with the narrow exception of DNSSEC-related records like RRSIG and NSEC. For example, developers occasionally attempt to place both an A record and a CNAME record on api.example.com, or attach a verification TXT record to a node that already hosts a CNAME. In this scenario, resolver behavior is undefined: some resolvers return only the CNAME, others return only the TXT, and some drop the response as a malformed zone entry.
Split-horizon DNS (internal versus external views) also triggers elusive troubleshooting errors. In split-horizon environments, cloud VPC internal resolvers return private IP addresses (e.g., 10.0.0.0/8) for internal hostnames, while public authoritative nameservers return public IPs or NXDOMAIN. When an engineer debugs a connection failure from their local workstation or a bastion host, dig queries the public authoritative infrastructure by default, masking the fact that an internal VPC resolver is serving a stale private view.
Finally, understand wildcard precedence rules. If a zone defines a wildcard record (*.example.com IN A 192.0.2.50), any explicit record defined at a specific subdomain completely supersedes the wildcard for that exact name. If api.example.com has an explicit TXT record but no A record, a query for api.example.com A will return NODATA (an empty answer with NOERROR) rather than falling back to the wildcard A record. Wildcards only match names that do not exist at all in the zone.
Building a Repeatable DNS Debugging Runbook for Your Team
Incident response teams need a clear, structured runbook to triage DNS failures rapidly without requiring deep protocol specialization under pressure. Standardize your team's incident workflow using this actionable execution runbook:
Step 1: Rapid Triage and Metric Verification
- Verify the failing FQDN, record type, and reporting client IP.
- Query a public recursive resolver:
dig @1.1.1.1 <FQDN> <TYPE>. Record thestatuscode. - Query with Checking Disabled:
dig @1.1.1.1 <FQDN> <TYPE> +cd. If this succeeds where the prior step failed, route immediately to DNSSEC remediation.
Step 2: Upstream Delegation Verification
- Query the parent TLD server for the zone's delegation:
dig @<tld-server> <zone> NS. - Ensure all returned NS records match your assigned authoritative nameservers.
- Confirm that in-bailiwick nameservers have valid glue records in the parent TLD response.
Step 3: Direct Authoritative Inspection
- Query every assigned authoritative nameserver independently using
+norecurse. - Verify that the
aaflag is set in the response headers. - Verify that the SOA serial numbers and record answers are consistently synchronized across all nameserver nodes.
Step 4: Operational Log Capture
- Capture the full output of
dig +trace <FQDN> <TYPE>anddelv @1.1.1.1 <FQDN> <TYPE>into the incident ticket. - Capture local resolver forwarding logs and system status (e.g.,
systemd-resolve --status).
Beyond manual runbooks, configure automated alerting on key DNS service-level indicators (SLIs). Monitor your authoritative nameserver query success rate, authoritative response latency, and global SERVFAIL percentages. Implement automated daily synthetic checks that validate your domain's DNSSEC chain from the root down to critical API endpoints, alerting your team days before an expiring RRSIG or mismatched DS record impacts production traffic.
DNSCove does not include dedicated DDoS scrubbing in v1. Runbook planning should account for upstream network protections. Furthermore, DNSCove serves standard authoritative records and does not offer GeoDNS, weighted, latency-based, or failover traffic steering in v1. Similarly, DNSCove does not offer AXFR zone transfer or secondary-DNS operation in v1. DNSCove does not expose a Route 53 wire-compatible API in v1; you manage DNS through DNSCove's own JSON API, console, and Terraform guides, and migrate off Route 53 with a one-step zone import. Review our comprehensive Route 53 migration guide to ensure zero downtime during zone transitions.
Conclusion: Debug From the Authoritative Source Outward
When DNS resolution breaks, guessing at root causes or flushing random caches wastes valuable time. The most reliable methodology for troubleshooting DNS resolution errors begins by querying the authoritative nameservers directly. Authoritative nameservers represent the single source of ground truth for your zone data.
Once you verify whether the authoritative nameservers return the expected records with the authoritative answer (aa) flag, work your way outward through the parent delegation hierarchy, the DNSSEC validation chain, and intermediate caching layers. By isolating each protocol layer with targeted dig and delv commands, you turn an opaque outage into an actionable, verifiable fix.
Frequently Asked Questions
Why does dig return SERVFAIL instead of NXDOMAIN for a name that does not exist?
When a requested name does not exist within a DNSSEC-signed zone, the authoritative nameserver must return an authenticated denial of existence using signed NSEC or NSEC3 records. If those NSEC/NSEC3 records are missing, expired, or have an invalid cryptographic signature, the validating recursive resolver cannot verify that the name truly does not exist. Because it cannot cryptographically prove the absence of the record, it assumes the response has been tampered with and returns SERVFAIL instead of NXDOMAIN.
How do I tell whether a DNS failure is at the resolver or the authoritative nameserver?
Query the authoritative nameserver directly using the +norecurse flag: dig +norecurse @ns1.yournameserver.com domain.com A. If the authoritative nameserver returns the expected record in the ANSWER SECTION with the aa (Authoritative Answer) flag present in the header, the authoritative layer is healthy. If a public resolver such as 1.1.1.1 or 8.8.8.8 fails to return that same answer, the breakdown is happening upstream at the resolver layer, typically due to parent delegation skew, firewall timeouts, or DNSSEC validation failures.
What causes a lame delegation and how do I fix it?
A lame delegation occurs when the parent registry (the TLD nameservers) delegates a domain to nameservers that are not configured to serve that domain, or are unreachable. Common causes include typos in the nameserver hostnames entered at your registrar, deleting a zone from your DNS provider before updating registrar delegation, or forgetting to configure child zones on secondary servers. To fix it, log in to your domain registrar's control panel and update your domain's delegation to point exclusively to active, correctly configured authoritative nameservers.
How long should I wait for a DNS change to propagate after lowering the TTL?
Lowering the TTL on a record only affects future lookups; it does not shorten the lifespan of records that have already been cached by resolvers. You must wait for the duration of the old, higher TTL before you can guarantee that all recursive resolvers worldwide have expired their stale cache entries and picked up the new, lower TTL. For example, if your original TTL was 86400 seconds (24 hours), you must wait 24 hours after publishing the reduced TTL before performing your production switchover.
Can a DNSSEC misconfiguration cause intermittent resolution failures for only some users?
Yes. Many consumer ISPs and older enterprise forwarders do not perform DNSSEC validation; they simply forward queries and return whatever the authoritative nameservers provide. Conversely, major public resolvers (such as Cloudflare, Google Public DNS, and Quad9) strictly validate DNSSEC signatures. If your zone suffers from an expired RRSIG, an orphaned DS record, or an invalid key digest, users resolving through validating resolvers will encounter persistent SERVFAIL errors, while users whose local resolvers ignore DNSSEC will resolve the domain without issue.
- No AWS account required
- Zero-downtime Route 53 cutover
- Apex ALIAS / ANAME to any target
- DNS as code — Terraform, CloudFormation
Straight answer: DNSSEC signing isn't available yet — it's on the roadmap. Everything else here works today. Authoritative nameservers: ns1.dnscove.com, ns2.dnscove.org.