The standby router that couldn't tell time
Two small VMs act as Tailscale subnet routers for the home LAN. Router 1
is primary and carries the 192.168.68.0/24 route into the tailnet.
Router 2 is standby: it advertises the same route, but the primary is the
one clients actually go through. If router 1 disappears, router 2 is
supposed to take over without anyone touching a config.
Both showed as online after the network cutover. That is the check I had been doing: is the standby connected, is the route advertised. It passed.
It was not fine.
The check that asked a second question
After the cutover I ran a health check on both routers and asked for more than up/down. Router 2 came back with two warnings that had nothing to do with the network swap:
Tailscale failed to fetch the DNS configuration of your device, with the first occurrence on 2026-09-15, before the cutoverSystem clock synchronized: no, with the clock about 1h28m behind
The standby had been in that state for almost two weeks, through a night where it was supposed to be the fallback path.
The DNS warning had a start date that predated the cutover, so the new hardware was not the cause. The old network had been hiding it just as well.
A clock problem that started at DNS
The instinct with a wrong clock is to look at NTP. The problem was one layer underneath it.
Router 2 had Tailscale DNS acceptance turned on. Router 1 did not. When
Tailscale manages DNS on a node, it writes the system resolver. On router
2, /etc/resolv.conf was a plain file Tailscale had written back in June,
pointing at Tailscale’s own resolver at 100.100.100.100. It was not a
symlink to any local resolver service, so nothing rewrote it later when
the machine’s real DNS changed.
That file is the whole chain:
- Name lookups went to
100.100.100.100, and on this node they failed. systemd-timesyncdcould not resolve the NTP pool, so it sent nothing (0 packets, not a rejected response).- With no time source, the clock drifted, which is why it read about 1h28m behind.
The clock warning was the visible symptom. DNS was the cause.
The fix, in order
Three steps, and the order mattered:
# 1. stop letting Tailscale manage this node's resolver
tailscale set --accept-dns=false
# 2. put the real resolver back, same two lines router 1 uses
# (search domain plus the Pi-hole VIP), keeping a backup
sudo cp /etc/resolv.conf /etc/resolv.conf.bak
sudo tee /etc/resolv.conf >/dev/null <<'EOF'
search <lan-domain>
nameserver 192.168.68.20
EOF
# 3. restart the time sync now that it can resolve a pool
sudo systemctl restart systemd-timesyncd
The part worth knowing: turning DNS acceptance off did not rewrite
/etc/resolv.conf. On this node the file stayed as Tailscale had left it
in June. --accept-dns=false stops Tailscale from managing the resolver
going forward, but restoring the file is a separate step. Without it, step
1 would have looked like a fix while name lookups stayed broken.
The result
After the resolver was restored, names resolved immediately, and
systemd-timesyncd reached a pool server and corrected the clock within
about a minute, settling to an offset near 1 ms. --accept-routes=false
stays as it was. That setting is correct on a node that advertises the LAN
rather than consuming routes.
The standby is now actually standby.
Why it mattered
A standby router is only a fallback if it can do the job when called. The failover test for this pair checks the route, and the route was there. What was missing was everything around it: if router 1 had failed, the takeover router would have had working routing and broken DNS, and its timestamps would have been an hour and a half off. The routing check would have said the failover worked.
The useful lesson is about which node gets checked. The primary is exercised constantly, so problems on it show up on their own. The standby sits idle and looks healthy, so it has to be checked on purpose, and the check has to look past “is it connected.” The other is that the clock was downstream. The timestamp warning was real, but fixing NTP first would have been fixing the symptom of a name resolution failure.