Today, 1.1.1.1 went dark for over two hours. The cause was not the attack everyone assumed.

Cloudflare's public DNS resolver suffered a global outage on July 14. Within minutes, monitoring tools showed Tata Communications (AS4755) announcing Cloudflare's 1.1.1.0/24 prefix — the textbook signature of a BGP hijack. That was not what took the service down.

WHAT ACTUALLY HAPPENED

The real root cause was a routine configuration change made over a year earlier, tied to a dormant Data Localization Suite deployment. It sat inert until engineers today attached a test location to that same service, triggering a global configuration refresh that withdrew 1.1.1.1's route announcements from Cloudflare's own production data centers worldwide.

The BGP hijack was real but incidental — a third party absorbing route space that Cloudflare had already vacated, visible only because Cloudflare's own announcements had gone silent first. Cloudflare detected the impact, declared an incident, and reverted the configuration; that rollback restored roughly 77% of traffic within twenty minutes.

WHY THE RED HERRING MATTERS

Every visible signal pointed to external attack: a foreign AS announcing your prefix is exactly what hijack detection tooling is built to flag, and it fired right on cue. The actual failure was internal — a change made over a year prior, dormant, then activated by an unrelated routine edit on an entirely different service. Attribution built on the first plausible signal, not the timeline, will point your incident response team at the wrong adversary while the real cause sits untouched in a config file.

WHAT TO DO NOW

TRACE CONFIGURATION CHANGES TO THEIR ORIGIN, NOT THEIR TRIGGER — the change that caused today's outage was made over a year before it activated. Audit trails need to survive that long.

DO NOT LET BGP HIJACK DETECTION BECOME YOUR DEFAULT ROOT CAUSE — verify your own route announcements before concluding a third party is the source.

TEST DORMANT CONFIGURATION PATHS, NOT JUST ACTIVE ONES — a change that "did nothing" for a year still shipped. Unused code paths in infrastructure config are latent risk, not inert risk.

KNOW YOUR ROLLBACK TIME BEFORE YOU NEED IT — Cloudflare recovered 77% of traffic in twenty minutes because reversion was fast and rehearsed. That is the metric that matters during an incident, not detection speed.

The internet's DNS layer went down today because of a change nobody was actively thinking about. That is the risk profile of modern infrastructure — not the attack you're watching for, but the one you forgot you shipped.

Does your incident response process default to "we were attacked" before ruling out "we did this to ourselves"?
Turn the analysis into a plan

The gap between knowing the risk and closing it is a purchase order and a weekend.

We specify, source and deploy the equipment that closes it — firewalls, segmentation, secure remote access — and we support it afterwards.