What a multi-agent AI taught me about finding bugs


I spend a lot of my time these days on two things that turn out to be the same thing: studying how vulnerabilities get found, and building an AI assistant that helps me find them. So when Anthropic put out research on using multi-agent systems for vulnerability research, it went straight to the top of my reading pile.

One number in it stuck with me. A group of agents working together surfaced hundreds of vulnerabilities, many times more than the same kind of agent running on its own. But the detail I actually can’t stop thinking about is a smaller one: how little the agents’ findings overlapped. Point several independent searchers at the same code and they don’t converge on one answer. They fan out. Each one wanders into a different corner and comes back holding something the others walked straight past.

That matches my own experience almost exactly, and it quietly overturns the mental model I started with.

The model I had wrong

When I first started automating security work, I treated it like a puzzle with a right answer. Get the prompt good enough, give the tool enough context, and it would “solve” the target. One smart pass. The bug either falls out or it doesn’t.

Real bug hunting isn’t shaped like that. A codebase or a network isn’t one puzzle, it’s a hundred small ones stacked together, and most of the work is noticing the ones you haven’t looked at yet. The failure mode is almost never “I looked at this and got it wrong.” It’s “I never looked at this at all.” You can be brilliant about the ten things you examined and still miss the eleventh, because you never turned it over.

Once you see the problem that way, the multi-agent result stops being surprising and starts being obvious. If the bottleneck is coverage, then the fix isn’t a cleverer single searcher. It’s more searchers who don’t all think alike.

Why difference beats brute force

Here’s the part that’s easy to get wrong. You don’t get more coverage just by running the same agent ten times. Run one identical searcher ten times and it tends to walk the same well-lit path on every pass. You pay ten times the cost for roughly one pass of results.

The value comes from making the searchers different. Give one a suspicious mind about authentication, another a nose for the ugly edge cases in input parsing, another that just asks “what happens if this runs twice at once.” They spread out because their instincts pull them in different directions. The low overlap in that research isn’t noise to be tidied away. It’s the whole point. Diversity of approach is what turns ten runs into ten genuinely different looks at the target.

I’ve felt this on my own setup. My earliest version was one assistant doing everything, and it was competent and repetitive in equal measure. It kept re-finding the same handful of things and calling it a day. The moment I split the work into narrower roles with different starting assumptions, the results stopped clustering. It found things the single version never went near — not because any one part got smarter, but because more of the target actually got looked at.

The catch: more findings is not more truth

There’s a trap on the other side of this, and I’d be lying if I said I hadn’t fallen into it. When you have a swarm of agents eagerly reporting things, you produce a lot of findings. Volume feels like progress. It usually isn’t.

A finding that nobody has confirmed is a claim, not a result. Ten agents producing forty plausible-sounding issues can be far worse than one careful pass producing three real ones, because now you have thirty-seven distractions dressed up as work. The more generation you add on the front, the more verification you need on the back — someone, or something, whose whole job is to try to disprove each finding rather than admire it.

So the shape I’ve landed on is two-sided. Spread out wide to find candidates, using searchers that deliberately disagree with each other. Then funnel hard, and treat every candidate as guilty of being false until it survives a real attempt to knock it down. Fan out to discover, squeeze down to believe. Skip the second half and you’ve just built a very fast machine for generating confident nonsense.

Why this stuck with me

I like this research because it puts a number on an instinct I’d been circling for a while. The win from pointing multiple minds at a hard problem doesn’t come from any of them being a genius. It comes from them being unalike, and from someone downstream being ruthless about what’s actually real.

That’s true of AI agents. It’s true of a bug bounty team. Honestly, it’s true of the study group I sit in at uni. The person who spots the thing everyone else missed is rarely the smartest one in the room. They’re the one who happened to be looking somewhere else. The trick, whether you’re wiring up agents or just working with other people, is to build a room full of people looking somewhere else on purpose — and then to check the answers before you believe them.