Bulletproof DNS


As someone with a deep interest in privacy and reliable systems - but not enough free time to deal with outages - I don't just want my home lab to work. I want it to be resilient, distributed, and enterprise-grade.

Recently, I decided to completely overhaul my DNS infrastructure. I run 4 distributed DNS servers on low cost 1L PCs running Proxmox as the base OS, split across two independent racks and subnets (10.5.0.x and 10.10.0.x). My goal was simple: create a load-balanced, high-availability system where I could patch, reboot, or lose an entire rack without my network noticing a blip.

Here's how I built a distributed DNS stack using AdGuard Home, DNSDist, Keepalived, and Ansible.

The Architecture

My design philosophy is simple: Layer 4 High Availability handles IP failover, Layer 7 Load Balancing handles DNS logic, and AdGuard Home does the actual filtering.

  1. Keepalived (Layer 4): Manages Virtual IPs (VIPs). If a node dies, the VIP floats to the next healthy node
  2. DNSDist (Layer 7): Sits on port 53, accepts traffic, checks upstream health (e.g., querying Cloudflare), and load balances queries to the backend
  3. AdGuard Home: The backend resolver. To avoid port conflicts, I moved this to port 5353
  4. AdGuardHome-Sync: Ensures that when I block a domain on my "Origin" server (10.0.0.70), it propagates to all 4 replicas instantly

The Topology

Step 1: Ansible Automation

I knew managing distributed servers manually would be a nightmare, so I used Ansible to orchestrate the entire configuration. This approach lets me reproduce the entire setup in minutes and keep all four servers in perfect sync.

The Inventory (inventory.ini)

I defined two groups, rack and telco, and assigned priorities for Keepalived. The highest priority gets the VIP master state.

[all:vars]
ansible_user=root
# Secure communication keys
dnsdist_console_key=ZRIyQ2Vz/X6J+PzJ9a0bK4lV8w3xN7dE5fG1hI2jK3m=
# Explicitly set Python to avoid warnings
ansible_python_interpreter=/usr/bin/python3

[rack]
adguard1 ansible_host=10.5.0.30 keepalived_priority=105 keepalived_ip=10.5.0.30
adguard2 ansible_host=10.5.0.31 keepalived_priority=104 keepalived_ip=10.5.0.31

[rack:vars]
vip_address=10.5.0.100/24
virtual_router_id=50

[telco]
adguard6        ansible_host=10.10.0.106 keepalived_priority=101 keepalived_ip=10.10.0.106
adguard-proxnas ansible_host=10.10.0.107 keepalived_priority=100 keepalived_ip=10.10.0.107

[telco:vars]
vip_address=10.10.0.100/24
virtual_router_id=51

[dns_cluster:children]
rack
telco

The Playbook (dns_cluster.yml)

The playbook moves AdGuard to port 5353, installs the necessary packages, and deploys my dynamic templates across all nodes.

---
- name: Deploy HA DNS Cluster (AdGuard + DNSDist + Keepalived)
  hosts: dns_cluster
  become: yes
  vars:
    adguard_config_path: "/opt/AdGuardHome/AdGuardHome.yaml"
    dnsdist_web_pass: "change_this_secure_password"
    dnsdist_api_key: "change_this_api_key"
    keepalived_interface: "eth0" # or ens18 for Proxmox
    keepalived_pass: "secure_auth_pass"

  tasks:
    - name: Ensure AdGuard Home is listening on port 5353
      replace:
        path: "{{ adguard_config_path }}"
        regexp: '^  port: 53$'
        replace: '  port: 5353'
        after: '^dns:'
      notify: Restart AdGuardHome

    - name: Install Dependencies
      apt:
        name:
          - dnsdist
          - keepalived
          - python3-pip
        state: present
        update_cache: yes

    - name: Deploy Configs
      template:
        src: "templates/{{ item }}.j2"
        dest: "/etc/{{ item | regex_replace('\\.conf.*', '') }}/{{ item | regex_replace('\\.j2$', '') }}"
      loop:
        - dnsdist.conf.j2
        - keepalived.conf.j2
      notify:
        - Restart DNSDist
        - Restart Keepalived

    - name: Enable Services
      systemd:
        name: "{{ item }}"
        enabled: yes
        state: started
      loop: [dnsdist, keepalived]

  handlers:
    - name: Restart AdGuardHome
      service: { name: AdGuardHome, state: restarted }
    - name: Restart DNSDist
      service: { name: dnsdist, state: restarted }
    - name: Restart Keepalived
      service: { name: keepalived, state: restarted }

Step 2: Configuration Templates

Keepalived Logic

I use Jinja2 templating to calculate the "Master" node automatically. The template finds the highest priority in each group and assigns it the MASTER state.

Template: keepalived.conf.j2

global_defs {
    router_id dns-lb-{{ inventory_hostname }}
}

{% set my_group = 'rack' if inventory_hostname in groups['rack'] else 'telco' %}
{% set max_prio = groups[my_group] | map('extract', hostvars, 'keepalived_priority') | map('int') | max %}

vrrp_instance VI_{{ my_group|upper }} {
    state {{ 'MASTER' if keepalived_priority|int == max_prio else 'BACKUP' }}
    interface {{ keepalived_interface }}
    virtual_router_id {{ virtual_router_id }}
    priority {{ keepalived_priority }}
    advert_int 1
    authentication {
        auth_type PASS
        auth_pass {{ keepalived_pass }}
    }
    unicast_src_ip {{ keepalived_ip }}
    unicast_peer {
{% for host in groups[my_group] %}
{% if host != inventory_hostname %}
        {{ hostvars[host]['keepalived_ip'] }}
{% endif %}
{% endfor %}
    }
    virtual_ipaddress {
        {{ vip_address }} dev {{ keepalived_interface }}
    }
}

DNSDist Logic (The Critical Part)

This is where I hit some snags during deployment that are worth sharing:

  1. ACLs: You must allow 127.0.0.0/8 explicitly. Without it, local health checks fail and dnsdist drops packets from itself
  2. Pools: I learned not to assign servers to named pools unless you have specific routing rules. By removing the pool= parameter, all servers go into the default pool, creating a massive load-balancing mesh

Template: dnsdist.conf.j2

-- 1. ACLs - CRITICAL: Allow localhost!
addACL('127.0.0.0/8')
addACL('::1/128')
addACL('10.0.0.0/8')
addACL('172.16.0.0/12')
addACL('192.168.0.0/16')

-- 2. Interfaces: Listen on all, including VIPs
setLocal('0.0.0.0:53')

-- 3. Webserver & API
webserver("0.0.0.0:8083")
setWebserverConfig({password="{{ dnsdist_web_pass }}", apiKey="{{ dnsdist_api_key }}"})

-- 4. Backends (All servers in the default pool)
{% for host in groups['rack'] + groups['telco'] %}
newServer({
    address="{{ hostvars[host]['keepalived_ip'] }}:5353",
    name="{{ host }}",
    checkName="cloudflare.com", -- L7 Health Check
    checkInterval=15
})
{% endfor %}

-- 5. Caching & Policy
pc = newPacketCache(10000)
getPool(""):setCache(pc)
setServerPolicy(leastOutstanding) -- Great for latency

-- 6. Console Access
setKey("{{ dnsdist_console_key }}")
controlSocket("127.0.0.1")

Step 3: Keeping it Synced

With 4 distributed servers, I'm definitely not logging into 4 different GUIs every time I want to block a domain. I use AdGuardHome-Sync running in Docker on my management node to propagate changes automatically.

Important note: Version 0.8.2 changed the config schema significantly. Here's my working config:

# adguardhome-sync.yaml
cron: "*/10 * * * *"
runOnStart: true
continueOnError: true
apiPort: 8080 

origin:
  url: "http://10.0.0.70" # My "Source of Truth"
  username: "admin"
  password: "password"

replicas:
  # Rack 1 Replicas
  - url: "http://10.5.0.30"
    username: "admin"
    password: "password"
  - url: "http://10.5.0.31"
    username: "admin"
    password: "password"

  # Telco Rack Replicas
  - url: "http://10.10.0.106"
    username: "admin"
    password: "password"
  - url: "http://10.10.0.107"
    username: "admin"
    password: "password"

features:
  generalSettings: true
  queryLogConfig: true
  statsConfig: true
  clientSettings: true
  services: true
  filters: true
  dhcpServerConfig: false # Don't sync DHCP!
  dnsAccessLists: true
  dnsServerConfig: true
  dnsRewrites: true

The Result

I now have a resilient DNS infrastructure where I can take down any single server - or an entire rack - and my devices (from my Rivian to my security cameras) keep resolving DNS without interruption.

Is it over-engineered? Maybe.

Is it bulletproof? Absolutely.

And more importantly, I can sleep at night knowing that my network won't go down because a single server decided to misbehave during a 2 AM automatic update.