BIT HAPPENSINFRA DIAGNOSTICS
⌕ FIND
EN
RUN DIAGNOSTICDIAG
MENU
DNS // TROUBLESHOOTING // SYSADMIN

How to troubleshoot DNS problems: a practical sysadmin workflow

A practical DNS troubleshooting workflow for sysadmins: isolate the scope, compare resolvers, query authoritative servers, check delegation and DNSSEC, and separate DNS failures from application problems.

Updated 2026-09-28Practical workflow
RUN DNS LOOKUP →

1. Define the symptom before changing DNS

"DNS is broken" can mean very different things: one workstation cannot resolve a name, an office resolver returns stale data, a public resolver returns SERVFAIL, the authoritative servers disagree, or DNS is correct and the real failure is HTTP, TLS, routing or firewall policy.

Start by recording the exact hostname, record type, client location, resolver in use, expected answer and observed answer. Do not flush caches immediately; cached evidence can help explain why only some users are affected.

RULE //

Change DNS only after you know which layer is returning the wrong answer.

2. Test the resolver the client is actually using

First reproduce the problem from the affected host. This tells you what the operating system and its configured recursive resolver currently see.

# Windows PowerShell
Resolve-DnsName example.com
Resolve-DnsName example.com -Type AAAA

# Windows classic
nslookup example.com

# Linux with systemd-resolved
resolvectl status
resolvectl query example.com

# Linux/macOS with dig
dig example.com A
dig example.com AAAA

Record the returned value, status and TTL. NXDOMAIN, SERVFAIL, timeout and a valid-but-wrong address point to different failure modes.

3. Compare independent recursive resolvers

If the configured resolver gives an unexpected answer, compare it with independent public resolvers. Different results usually indicate cache state, filtering, split-horizon DNS or a resolver-specific problem rather than an authoritative-zone change.

dig @1.1.1.1 example.com A
dig @8.8.8.8 example.com A
dig @9.9.9.9 example.com A

Do not treat a public resolver as the source of truth. The next step is to bypass recursive caches and inspect the authoritative path.

4. Find the authoritative nameservers and query them directly

Identify the NS records, then query each authoritative server directly for the record that is failing.

dig NS example.com
dig @ns1.example-dns.net example.com A
dig @ns2.example-dns.net example.com A

Every authoritative server for the same zone should provide a coherent view. If one server returns old data, times out or has a different SOA serial, the problem is likely in zone distribution or authoritative infrastructure.

The DNS Lookup tool is useful for quickly inspecting the record set, while Domain Health provides a broader domain-level view.

5. Check delegation, glue and DNSSEC when answers look inconsistent

A correct zone can still fail if the parent zone delegates to the wrong nameservers, required glue is missing or stale, or DNSSEC validation breaks between the parent and child zones.

dig +trace example.com
dig DS example.com
dig DNSKEY example.com +dnssec
dig example.com A +dnssec

With DNSSEC, a common operational pattern is that non-validating tests appear to work while validating recursive resolvers return SERVFAIL. When DNSSEC is involved, inspect the DS record at the parent, the DNSKEY set in the child zone and the signature chain before changing application infrastructure.

6. Use TTLs to explain propagation instead of guessing

DNS changes do not "propagate" as a single global event. Recursive resolvers cache answers until their TTL expires. If an old record had a long TTL before the change, some clients may legitimately keep the previous answer until that cached value ages out.

Compare the answer and remaining TTL across resolvers. Flush a local cache only after you have captured the evidence you need.

# Windows
ipconfig /displaydns
ipconfig /flushdns

# systemd-resolved
resolvectl flush-caches

7. Prove whether DNS is actually the failing layer

Once the hostname resolves to the expected IP, continue upward. A correct DNS answer does not prove that TCP, TLS or HTTP is healthy.

curl -I https://example.com
curl -vk https://example.com/

# Test the application against a specific IP while preserving the hostname/SNI
curl --resolve example.com:443:203.0.113.10 https://example.com/

If --resolve works against the intended IP while normal resolution sends the client elsewhere, DNS is still implicated. If both fail identically, investigate routing, firewall/WAF, listener configuration, certificate/SNI or the application itself. For route-level investigation, use Traceroute.

Common failure patterns

  • Only one workstation fails: local cache, VPN, hosts file, endpoint DNS policy or local resolver configuration.
  • Only one office/site fails: internal resolver, forwarder, conditional zone, split DNS or network policy.
  • Public resolvers disagree temporarily: cache/TTL state or resolver-specific behavior.
  • Authoritative servers disagree: zone replication or provider configuration problem.
  • Validating resolvers return SERVFAIL: investigate DNSSEC chain and signatures.
  • DNS returns the expected IP but the service is down: move to TCP, TLS, HTTP, firewall and application diagnostics.

8. A repeatable DNS troubleshooting checklist

  1. Write down the exact hostname, record type, expected value and affected users.
  2. Query from an affected client using its configured resolver.
  3. Compare at least two independent recursive resolvers.
  4. Discover the authoritative nameservers.
  5. Query each authoritative server directly.
  6. Inspect delegation and DNSSEC when the chain is suspicious.
  7. Use TTL values to explain stale answers.
  8. Once DNS is correct, test TCP/TLS/HTTP separately.
  9. Change only the layer that produced the bad evidence, then repeat the same test.

This workflow keeps DNS troubleshooting evidence-driven: client → recursive resolver → authoritative DNS → delegation/DNSSEC → application layer.

NEXT STEP //

Run the same checks against the affected domain, keep the evidence from each layer, and change only the component that actually failed.

RUN DNS LOOKUP →MORE GUIDES →