Bulletproof DNS
As someone with a deep interest in privacy and reliable systems - but not enough free time to deal with outages - I don't just want my home lab to work. I want it to be resilient, distributed, and enterprise-grade.
Recently, I decided to completely overhaul my DNS infrastructure. I run 4 distributed DNS servers on low cost 1L PCs running Proxmox as the base OS, split across two independent racks and subnets (10.5.0.x and 10.10.0.x). My goal was simple: create a load-balanced, high-availability system where I could patch, reboot, or lose an entire rack without my network noticing a blip.
Here's how I built a distributed DNS stack using AdGuard Home, DNSDist, Keepalived, and Ansible.
The Architecture
My design philosophy is simple: Layer 4 High Availability handles IP failover, Layer 7 Load Balancing handles DNS logic, and AdGuard Home does the actual filtering.
- Keepalived (Layer 4): Manages Virtual IPs (VIPs). If a node dies, the VIP floats to the next healthy node
- DNSDist (Layer 7): Sits on port 53, accepts traffic, checks upstream health (e.g., querying Cloudflare), and load balances queries to the backend
- AdGuard Home: The backend resolver. To avoid port conflicts, I moved this to port
5353 - AdGuardHome-Sync: Ensures that when I block a domain on my "Origin" server (
10.0.0.70), it propagates to all 4 replicas instantly
The Topology
- Rack 1 (10.5.0.0/24): 2 servers (
adguard1,adguard2) with VIP:10.5.0.100 - Telco Rack (10.10.0.0/24): 2 servers (
adguard3,adguard4) with VIP:10.10.0.100
Step 1: Ansible Automation
I knew managing distributed servers manually would be a nightmare, so I used Ansible to orchestrate the entire configuration. This approach lets me reproduce the entire setup in minutes and keep all four servers in perfect sync.
The Inventory (inventory.ini)
I defined two groups, rack and telco, and assigned priorities for Keepalived. The highest priority gets the VIP master state.
[all:vars]
ansible_user=root
# Secure communication keys
dnsdist_console_key=ZRIyQ2Vz/X6J+PzJ9a0bK4lV8w3xN7dE5fG1hI2jK3m=
# Explicitly set Python to avoid warnings
ansible_python_interpreter=/usr/bin/python3
[rack]
adguard1 ansible_host=10.5.0.30 keepalived_priority=105 keepalived_ip=10.5.0.30
adguard2 ansible_host=10.5.0.31 keepalived_priority=104 keepalived_ip=10.5.0.31
[rack:vars]
vip_address=10.5.0.100/24
virtual_router_id=50
[telco]
adguard6 ansible_host=10.10.0.106 keepalived_priority=101 keepalived_ip=10.10.0.106
adguard-proxnas ansible_host=10.10.0.107 keepalived_priority=100 keepalived_ip=10.10.0.107
[telco:vars]
vip_address=10.10.0.100/24
virtual_router_id=51
[dns_cluster:children]
rack
telcoThe Playbook (dns_cluster.yml)
The playbook moves AdGuard to port 5353, installs the necessary packages, and deploys my dynamic templates across all nodes.
---
- name: Deploy HA DNS Cluster (AdGuard + DNSDist + Keepalived)
hosts: dns_cluster
become: yes
vars:
adguard_config_path: "/opt/AdGuardHome/AdGuardHome.yaml"
dnsdist_web_pass: "change_this_secure_password"
dnsdist_api_key: "change_this_api_key"
keepalived_interface: "eth0" # or ens18 for Proxmox
keepalived_pass: "secure_auth_pass"
tasks:
- name: Ensure AdGuard Home is listening on port 5353
replace:
path: "{{ adguard_config_path }}"
regexp: '^ port: 53$'
replace: ' port: 5353'
after: '^dns:'
notify: Restart AdGuardHome
- name: Install Dependencies
apt:
name:
- dnsdist
- keepalived
- python3-pip
state: present
update_cache: yes
- name: Deploy Configs
template:
src: "templates/{{ item }}.j2"
dest: "/etc/{{ item | regex_replace('\\.conf.*', '') }}/{{ item | regex_replace('\\.j2$', '') }}"
loop:
- dnsdist.conf.j2
- keepalived.conf.j2
notify:
- Restart DNSDist
- Restart Keepalived
- name: Enable Services
systemd:
name: "{{ item }}"
enabled: yes
state: started
loop: [dnsdist, keepalived]
handlers:
- name: Restart AdGuardHome
service: { name: AdGuardHome, state: restarted }
- name: Restart DNSDist
service: { name: dnsdist, state: restarted }
- name: Restart Keepalived
service: { name: keepalived, state: restarted }Step 2: Configuration Templates
Keepalived Logic
I use Jinja2 templating to calculate the "Master" node automatically. The template finds the highest priority in each group and assigns it the MASTER state.
Template: keepalived.conf.j2
global_defs {
router_id dns-lb-{{ inventory_hostname }}
}
{% set my_group = 'rack' if inventory_hostname in groups['rack'] else 'telco' %}
{% set max_prio = groups[my_group] | map('extract', hostvars, 'keepalived_priority') | map('int') | max %}
vrrp_instance VI_{{ my_group|upper }} {
state {{ 'MASTER' if keepalived_priority|int == max_prio else 'BACKUP' }}
interface {{ keepalived_interface }}
virtual_router_id {{ virtual_router_id }}
priority {{ keepalived_priority }}
advert_int 1
authentication {
auth_type PASS
auth_pass {{ keepalived_pass }}
}
unicast_src_ip {{ keepalived_ip }}
unicast_peer {
{% for host in groups[my_group] %}
{% if host != inventory_hostname %}
{{ hostvars[host]['keepalived_ip'] }}
{% endif %}
{% endfor %}
}
virtual_ipaddress {
{{ vip_address }} dev {{ keepalived_interface }}
}
}DNSDist Logic (The Critical Part)
This is where I hit some snags during deployment that are worth sharing:
- ACLs: You must allow
127.0.0.0/8explicitly. Without it, local health checks fail anddnsdistdrops packets from itself - Pools: I learned not to assign servers to named pools unless you have specific routing rules. By removing the
pool=parameter, all servers go into the default pool, creating a massive load-balancing mesh
Template: dnsdist.conf.j2
-- 1. ACLs - CRITICAL: Allow localhost!
addACL('127.0.0.0/8')
addACL('::1/128')
addACL('10.0.0.0/8')
addACL('172.16.0.0/12')
addACL('192.168.0.0/16')
-- 2. Interfaces: Listen on all, including VIPs
setLocal('0.0.0.0:53')
-- 3. Webserver & API
webserver("0.0.0.0:8083")
setWebserverConfig({password="{{ dnsdist_web_pass }}", apiKey="{{ dnsdist_api_key }}"})
-- 4. Backends (All servers in the default pool)
{% for host in groups['rack'] + groups['telco'] %}
newServer({
address="{{ hostvars[host]['keepalived_ip'] }}:5353",
name="{{ host }}",
checkName="cloudflare.com", -- L7 Health Check
checkInterval=15
})
{% endfor %}
-- 5. Caching & Policy
pc = newPacketCache(10000)
getPool(""):setCache(pc)
setServerPolicy(leastOutstanding) -- Great for latency
-- 6. Console Access
setKey("{{ dnsdist_console_key }}")
controlSocket("127.0.0.1")Step 3: Keeping it Synced
With 4 distributed servers, I'm definitely not logging into 4 different GUIs every time I want to block a domain. I use AdGuardHome-Sync running in Docker on my management node to propagate changes automatically.
Important note: Version 0.8.2 changed the config schema significantly. Here's my working config:
# adguardhome-sync.yaml
cron: "*/10 * * * *"
runOnStart: true
continueOnError: true
apiPort: 8080
origin:
url: "http://10.0.0.70" # My "Source of Truth"
username: "admin"
password: "password"
replicas:
# Rack 1 Replicas
- url: "http://10.5.0.30"
username: "admin"
password: "password"
- url: "http://10.5.0.31"
username: "admin"
password: "password"
# Telco Rack Replicas
- url: "http://10.10.0.106"
username: "admin"
password: "password"
- url: "http://10.10.0.107"
username: "admin"
password: "password"
features:
generalSettings: true
queryLogConfig: true
statsConfig: true
clientSettings: true
services: true
filters: true
dhcpServerConfig: false # Don't sync DHCP!
dnsAccessLists: true
dnsServerConfig: true
dnsRewrites: trueThe Result
I now have a resilient DNS infrastructure where I can take down any single server - or an entire rack - and my devices (from my Rivian to my security cameras) keep resolving DNS without interruption.
Is it over-engineered? Maybe.
Is it bulletproof? Absolutely.
And more importantly, I can sleep at night knowing that my network won't go down because a single server decided to misbehave during a 2 AM automatic update.