TL;DR
I had 63 flaky tests in a Node.js monorepo that made every CI run a coin flip. Over one week I pointed Claude Code at them with a reproduction harness instead of a "please fix this" prompt, and it fixed 58 of them for real, quarantined 5, and taught me a lot about what AI coding agents are actually good at. Spoiler: the bottleneck was never the fixing. It was the diagnosis. 🚀
The Problem
Our monorepo had around 4,200 tests split across 14 packages. Vitest for unit tests, Playwright for the browser stuff, a handful of integration tests hitting a local Postgres. On paper, coverage was fine.
In practice, main was red about 30% of the time for no reason. Someone would push a one-line docs change, CI would fail, they'd hit "re-run failed jobs," and it would go green. We had a Slack emoji for it. 🎲
The costs were real:
Merge latency: every PR needed on average 1.7 CI runs to land. At 41 minutes per run, that's a lot of dead time.
Trust erosion: people
Discussion
Be the first to comment
Add your perspective to get the discussion started.