Routing and WAN
Routing and WAN
Two ISPs, OSPF with BFD, a canary route, and a CHR hot standby. The core does not originate default; it installs whichever ISP is cheaper and currently adjacent.

Dual ISP
| Path | Box | WAN | Transit to core | OSPF cost | BFD |
|---|---|---|---|---|---|
| Primary | CCR2004 Quickline 10.9.8.200 |
10G on CRS317 | vlan33 192.168.33.0/24 (core sfp-sfpplus4 ↔ edge sfp-sfpplus2) |
10 | yes |
| Backup | CCR2004 Wingo 10.9.8.216 |
1G on ether1 | vlan32 192.168.32.0/24 (core sfp-sfpplus1 ↔ edge sfp-sfpplus2) |
20 | yes |
Dashboard /wan names the active path (quickline expected) and whether the Wingo canary is alive. Fail-over is an audited drill (docs/wan-failover-drill-dash06.md in the dashboard repo), not a DHCP-timeout flip.
Wingo also has a VPN fallback dst-nat of UDP/13231 toward apu01 (wg-home). That path is configured; the handshake drill is still awaiting a bodo window.
VRRP standby (FAB-13)
core-sb is a RouterOS CHR on hp02 (10.9.8.251).
- VRRP backup, priority 100, on SVIs 7 / 35 / 50 / 58 / 59 / 62 / 90 / 99.
- Real address
.3, VIP.1(the address clients already use). - Own OSPF to both edges: vlan34 cost 30 (Quickline copy), vlan36 cost 40 (Wingo copy).
- DHCP servers stay disabled until VRRP master.
- Config is generated from the live core (
scripts/gen_core_standby.py). Drift is the standby tile on the dashboard. - Trunk is hp02
eno5np0→ CRS326-Csfp-sfpplus4. That port is DHCP-snooping trusted so a real core failure can still offer leases.
vlan53 (DNS) and vlan81 are not on the standby. The APUs stay pinned to the CCR2116.
flowchart LR
QL[Quickline edge] -->|"vlan33 cost 10 + BFD"| CORE[CCR2116]
WI[Wingo edge] -->|"vlan32 cost 20 + BFD"| CORE
QL -->|"vlan34 cost 30"| SB[core-sb CHR]
WI -->|"vlan36 cost 40"| SB
CORE <-->|"VRRP VIP .1"| SB
CORE -->|"vlan53 + BFD"| APU[apu01 / apu02]
APU -->|"10.53.53.53/32 stub"| CORE
DNS in the routing picture
Clients send DNS to 10.53.53.53. Each APU puts that address on lo and FRR advertises it as an OSPF stub. The core ECMP-hashes per source/destination pair. One probe source only ever sees one arm — that is why /dns probes from both 172.16.62.253 and 10.9.8.253.
If an APU cannot serve DNS it withdraws the address from lo (anycast healthcheck). OSPF Full with an empty kernel FIB is not enough; that was the Sep 2026 blackhole. Full story: DNS APUs — Anycast self-heal.
The core also intercepts rogue DNS (dst-nat of UDP/TCP 53 not already aimed at 10.53.53.53, from sources not on DNS-SERVERS) and sends it to the anycast.
Management-plane routing
vlan89 is a different router: CCR2004-16G 10.9.8.1, own WAN (213.221.211.27) and WireGuard (10.9.9.1). Every fabric device joins it with an unbridged L3 port. That island surviving when production RSTP or the CCR2116 is down is the point.
OOB recovery path: WireGuard → 10.9.8.1. Keep that reachable in every change plan.
hp02 talks RouterOS REST only from 10.9.8.253 (and its own vlan62 address is not used to reach routers). hp04 has a public /28 address on vlan448 and a vlan62 NIC; it must not forward between them. Agents on hp04 jump through hp02 for every fabric command.
What /wan and /overview should look like when healthy
From the 2026-09-16 dashboard snapshot used to write this book:
- devices 21 / 21 up, watched uplinks up
- OSPF 4 / 4 Full, BFD 6 / 6 up
- active default via Quickline, canary active
- pathproof 16 / 16 green
- DNS probes 9 / 9, anycast 3 / 3, cache hit ~99 %
No comments to display
No comments to display