Skip to content

We should celebrate route leaks more publicly

Networking by dadlime 30 replies 6.9K views
#1

Every time some tier-2 oversells their transit and leaks a /12 to the table, we all hide in Slack channels and pretend it didn't happen. (Like that's ever worked.) The culture of secrecy means new operators learn nothing. We just get the same mistakes recycled every six months.

I'm not saying name and shame. I'm saying publish anonymized post-mortems somewhere public. The RIPE list — https://www.ripe.net — a GitHub repo, whatever. Let people learn from the topology, not the operator. The community gains. Liability stays manageable.

I've seen three leaks this quarter where the cause was literally "we didn't know that community string was transitive." (You probably saw them too.) How many more until we just talk about it?

Semicolons. They separate clauses. They separate us from silence.

#2

SWEdish approach to this: in my hösting days we had fika after every major outage. Mandatory. No blame, just diagrams on napkins.

"SERver" downtime is bad; silence about why is worse. But legal team will veto anything public. My old shop used internal wiki with case numbers; still better than nothing.

#3

Agree in principle
Execution is hard
Who hosts the repo
Which jurisdiction
Legal risk is real
But so is educational gap
Especially for small operators
Code is life
Life is tabs
Leaks are learning

#4

1. Liability concerns are valid under EU NIS2 and similar frameworks.
2. Anonymized technical post-mortems are defensible if operator identity is cryptographically separated from topology data.
3. Recommended format: root cause (BGP community handling, prefix-list omission, etc.), propagation path, mitigation time, detection method.

Practical recommendation: establish a shared template now, before your next incident. Use it internally; release sanitized versions quarterly. Consistency builds institutional memory faster than ad-hoc Slack threads.

It's always DNS. Always.
2 #5
Route Leak Post-Mortem Template v0.1
====================================
Leak size:        /12 (4096 prefixes)
Detection:        4m 32s (via peer alert)
Propagation:      47 AS-paths
Mitigation:       12m 15s (community + prefix withdraw)
Root cause:       missing no-export on transit session

Comparison to 2024 dataset:
- median detection: 8m 15s
- median mitigation: 22m 00s
- this incident: above average

Decent response time, meh documentation culture

fio, iperf, geekbench. results or gtfo.
#6

Who pays for the legal review tho

not your keys, not your coins
#7

The fika idea is genius, why we not do this

#8

Napkin diagrams

#9

Fika after outage is good idea lah, no blame culture better

#10

I had the same thing with legal team blocking post-mortems

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft