How to troubleshoot DNS problems: a practical sysadmin workflow
A practical DNS troubleshooting workflow for sysadmins: isolate the scope, compare resolvers, query authoritative servers, check delegation and DNSSEC, and separate DNS failures from application problems.
RUN DNS LOOKUP →1. Define the symptom before changing DNS
"DNS is broken" can mean very different things: one workstation cannot resolve a name, an office resolver returns stale data, a public resolver returns SERVFAIL, the authoritative servers disagree, or DNS is correct and the real failure is HTTP, TLS, routing or firewall policy.
Start by recording the exact hostname, record type, client location, resolver in use, expected answer and observed answer. Do not flush caches immediately; cached evidence can help explain why only some users are affected.
Change DNS only after you know which layer is returning the wrong answer.
2. Test the resolver the client is actually using
First reproduce the problem from the affected host. This tells you what the operating system and its configured recursive resolver currently see.
# Windows PowerShell
Resolve-DnsName example.com
Resolve-DnsName example.com -Type AAAA
# Windows classic
nslookup example.com
# Linux with systemd-resolved
resolvectl status
resolvectl query example.com
# Linux/macOS with dig
dig example.com A
dig example.com AAAARecord the returned value, status and TTL. NXDOMAIN, SERVFAIL, timeout and a valid-but-wrong address point to different failure modes.
3. Compare independent recursive resolvers
If the configured resolver gives an unexpected answer, compare it with independent public resolvers. Different results usually indicate cache state, filtering, split-horizon DNS or a resolver-specific problem rather than an authoritative-zone change.
dig @1.1.1.1 example.com A
dig @8.8.8.8 example.com A
dig @9.9.9.9 example.com ADo not treat a public resolver as the source of truth. The next step is to bypass recursive caches and inspect the authoritative path.
5. Check delegation, glue and DNSSEC when answers look inconsistent
A correct zone can still fail if the parent zone delegates to the wrong nameservers, required glue is missing or stale, or DNSSEC validation breaks between the parent and child zones.
dig +trace example.com
dig DS example.com
dig DNSKEY example.com +dnssec
dig example.com A +dnssecWith DNSSEC, a common operational pattern is that non-validating tests appear to work while validating recursive resolvers return SERVFAIL. When DNSSEC is involved, inspect the DS record at the parent, the DNSKEY set in the child zone and the signature chain before changing application infrastructure.
6. Use TTLs to explain propagation instead of guessing
DNS changes do not "propagate" as a single global event. Recursive resolvers cache answers until their TTL expires. If an old record had a long TTL before the change, some clients may legitimately keep the previous answer until that cached value ages out.
Compare the answer and remaining TTL across resolvers. Flush a local cache only after you have captured the evidence you need.
# Windows
ipconfig /displaydns
ipconfig /flushdns
# systemd-resolved
resolvectl flush-caches7. Prove whether DNS is actually the failing layer
Once the hostname resolves to the expected IP, continue upward. A correct DNS answer does not prove that TCP, TLS or HTTP is healthy.
curl -I https://example.com
curl -vk https://example.com/
# Test the application against a specific IP while preserving the hostname/SNI
curl --resolve example.com:443:203.0.113.10 https://example.com/If --resolve works against the intended IP while normal resolution sends the client elsewhere, DNS is still implicated. If both fail identically, investigate routing, firewall/WAF, listener configuration, certificate/SNI or the application itself. For route-level investigation, use Traceroute.
Common failure patterns
- Only one workstation fails: local cache, VPN, hosts file, endpoint DNS policy or local resolver configuration.
- Only one office/site fails: internal resolver, forwarder, conditional zone, split DNS or network policy.
- Public resolvers disagree temporarily: cache/TTL state or resolver-specific behavior.
- Authoritative servers disagree: zone replication or provider configuration problem.
- Validating resolvers return SERVFAIL: investigate DNSSEC chain and signatures.
- DNS returns the expected IP but the service is down: move to TCP, TLS, HTTP, firewall and application diagnostics.
8. A repeatable DNS troubleshooting checklist
- Write down the exact hostname, record type, expected value and affected users.
- Query from an affected client using its configured resolver.
- Compare at least two independent recursive resolvers.
- Discover the authoritative nameservers.
- Query each authoritative server directly.
- Inspect delegation and DNSSEC when the chain is suspicious.
- Use TTL values to explain stale answers.
- Once DNS is correct, test TCP/TLS/HTTP separately.
- Change only the layer that produced the bad evidence, then repeat the same test.
This workflow keeps DNS troubleshooting evidence-driven: client → recursive resolver → authoritative DNS → delegation/DNSSEC → application layer.