Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.
How does Amazon stay up during Black Friday when a normal server would be on fire by 9am? Why do banks rarely lose a transaction even when a data center loses power mid-transfer?
The honest answer is not "they don't have failures." Everyone has failures.
Networks partition, disks die, servers get evicted by a bored Kubernetes scheduler for no reason anyone can explain.
Reliability isn't about preventing that.
It's about the system still doing the right thing while it's happening.
There are seven ideas that keep showing up whenever you dig into how large systems actually stay reliable.
I went through and a couple of them connect straight back to Kademlia and XOR distance, which I wrote about a while back.
Small world.
Kademlia: A
Discussion
Get the discussion rolling
A single comment can start something great.