aboutrssgithubx
babysitting-an-army-of-monkeys-2-planet-of-the-agents

Babysitting an Army of Monkeys 2: Planet of the Agents

What happens when Charlie Miller's dumb, honest monkeys learn to reason, and to lie. Babysitting used to mean counting crashes; with LLM agents it means refereeing liars and rebuilding the free oracle the kernel used to give you.

fuzzingagentsllmoraclesverification

what happens when Charlie Miller's dumb, honest monkeys learn to reason, and to lie

TL;DR In 2010, Charlie Miller fuzzed four products with five lines of Python and let a swarm of dumb monkeys pound the keyboard. It worked beautifully, and it still works, because the kernel hands you ground truth for free: a crash is a crash. Swap the monkeys for LLM agents and the ground shifts underneath you. They're expensive, they forget, they're wildly inconsistent, and worst of all they lie. One study caught a model declaring an attack finished before it had actually broken in, in half its runs. The free oracle is gone. And once you see that, every serious agentic-security system in 2026 turns out to be doing the same secret thing: rebuilding the oracle the monkey never needed. This is a long ramble about why babysitting stopped meaning counting crashes and started meaning refereeing liars. πŸ’


Back in 2010, Charlie Miller got up at CanSecWest and gave a talk with maybe the best title our field has ever produced: "Babysitting an Army of Monkeys." The pitch was beautiful in its simplicity: fuzz four products with about five lines of Python. Take some valid files, mutate the hell out of them, fling millions at a target, and go get coffee while your monkeys pound the keyboard. He turned ~1500 seed files into ~3 million mutated ones, threw them at Adobe Reader and Preview, and watched stuff fall over. Around five percent of what he fed Preview crashed it.

Here are the five lines. This is the entire engine:

# Charlie Miller, "Babysitting an Army of Monkeys" (2010)
numwrites = random.randrange(math.ceil((float(len(buf)) / FuzzFactor))) + 1
for j in range(numwrites):
    rbyte = random.randrange(256)
    rn = random.randrange(len(buf))
    buf[rn] = "%c" % (rbyte)

That's the whole monkey. Read a valid file into buf, flip a handful of random bytes to random values, write it back out, throw it at the target, watch for a crash. No grammar, no model, no coverage feedback, nothing clever.

And I'm not being sentimental when I say a variation of those exact lines has sat at the core of most fuzzers I've ever built. Apple's ICC profile parsing, PDF parsers, font parsers, and the meaner targets too: basebands, browser fuzzers. Different harnesses, different seed corpora, different piles of crash-triage duct tape bolted around the edges, but the beating heart was almost always Charlie's byte-flipper. That cute little loop has found me more real bugs than most of the clever things I ever wrote.

inspired by the old infinite-monkey-theorem gag: sit enough monkeys at enough typewriters and eventually one bangs out Shakespeare, you can sit enough monkeys at enough file formats and eventually one bangs out a segfault.

Quick note before I start poking at what changed: this is a sequel. Miller's talk was the jumping-off point, and dumb fuzzing still finds bugs every single day, sixteen years later. I'm taking his exact setup, keeping what made it work, and checking things underneath it to see what breaks.

Because something did change, and it's the whole reason for this post. Sixteen years on, the monkeys have evolved. Somewhere along the way they stopped flailing and started thinking: they read the codebase now, reason about it, write you a tidy report. And like every good Planet of the Apes sequel warns you, the moment the monkeys learn to talk is the moment they learn to lie. πŸ’

Everyone I know is now trying to babysit this new, smarter, chattier troop: XBOW, RunSybil, Hacktron, Snowball, the Ralph-loopers, Devin, a pile of new academic benchmarks, and in a small way, me.

So here's the question I can't stop poking at: what does "Babysitting an Army of Monkeys" turn into when the monkeys are AI agents?

Short version: every property that made the monkeys work depends on conditions that agents quietly take away. And the way they take them away tells you exactly what a modern security harness is really for. Let's go.

The old monkeys were perfect because they were dumb

Full disclosure: this was one of the earliest talks that pulled me properly into fuzzing, so I've had years to project my own reading onto it. Take what follows as my takeaway. But the bit that stuck with me wasn't "fuzzing finds bugs." It was this: fuzzing, to me, isn't about creating test cases. It's about filtering. You don't lovingly craft inputs. You generate an obscene number of them for free and let the crashes filter themselves out of the noise. Generating is trivial. Triage is the bottleneck. Fuzzing gave both exploration and exploitation for free.

Here's the loop, roughly:

        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚  seed files  β”‚
        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
               β”‚ mutate  (dumb Β· cheap Β· millions)
               β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        the KERNEL decides,
   β”Œβ”€β”€β”€β–Ίβ”‚  run target  β”‚        not the monkey
   β”‚    β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
   β”‚ next      β”‚
   β”‚ input β”Œβ”€β”€β”€β–Όβ”€β”€β”€β”€β”
   └───────│ crash? │──── no
           β””β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
               β”‚ yes   ◄── FREE, HONEST GROUND TRUTH
               β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚ triage (you) β”‚   ◄── the only bottleneck
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The monkeys had three properties that made this whole thing work, even if nobody bothered to put them on a slide:

They were cheap. A run cost microseconds. Who cares if 99.99% are useless? Make more.

They were dumb. No memory, no plan. Input number two-million has no idea what input one-million did. And that's fine: you didn't want them thinking, you wanted flailing, fast.

And they were honest. This is the one that mattered most and got noticed least. A segfault is a segfault. The OS handed you ground truth for free. The monkey physically could not lie about crashing, because it wasn't the one deciding. Your "babysitting" was just watching a pile of honest crashes roll in and picking out the spicy ones.

Cheap, dumb, honest. Hold onto those three. Now let's go hire some evolved monkeys, the "ai agents" and watch all three properties catch fire, not because the monkeys were doing it wrong, but because the agents are standing on different ground. πŸ”₯

The new monkeys are expensive, smart, and little liars

Swap the mutation loop for a fleet of LLM agents auditing code, and every single property flips. Let me take them one at a time, because each inversion is its own headache and the industry has already stubbed its toes on all of them.

Cheap β†’ expensive πŸ’Έ

A monkey was microseconds. An agent run is real dollars and real minutes.

The numbers are genuinely startling once you go looking. Anthropic's own writeup on their multi-agent research system notes it burns something like fifteen times the tokens of a normal chat, and that token usage alone explained the lion's share of the performance variance; the system is good because it's allowed to spend. XBOW, in their Alloys research, cut their solver loops off at 80 iterations in their experiments, because past that it's cheaper to start a fresh agent than to keep paying the current one. Other agentic companies budget their scanner per repo with hard task caps and worker pools; their worst single run took over fourteen hours. Devin started using a new billing unit (the "ACU") for how much agent-compute a task eats.

You are not throwing three million of these at a wall. And this is the important part: Miller's "generate infinite garbage and filter" wasn't naive. It was pure genius for a world where a run costs nothing. Flip the cost model, and the exact same strategy becomes unaffordable because the ground moved. You're back to the oldest problem there is: spend every run like it counts.

Dumb β†’ smart (but with the object permanence of a goldfish) 🐠

Great, it reasons now! It holds hypotheses, reads the code, adapts. Except it fills its context window after skimming one corner of a real repo, and then quietly forgets the bug it found this morning during context compaction.

This one's so universal it spawned a whole technique. Geoffrey Huntley's Ralph loop, a coding agent run in a dumb bash while loop, works specifically because it throws away context every iteration. Each turn is a fresh agent that reads its state off the filesystem (a TODO file, git history), does one thing, commits, and dies. Fresh context isn't a side effect, it's the entire point, because models measurably rot as the window fills. XBOW restart their solvers for the same reason. Cloudflare in their recent article keep each agent below a quarter of its context window and externalize everything else to a database.

I've messed with this pattern myself. One of my throwaway projects is a little Ralph variant called ChiefWiggum that runs the same fresh-context loop but points it at hunting vulnerabilities instead of writing code. Named because Chief Wiggum is Ralph's dad: the loop's the idiot son, somebody's gotta stand behind it. This has yielded me some goooood chunk of bounties.

And here's a quieter horror the smart monkey introduced: it isn't even repeatable. One 2026 study ran the same model, against the same target, with the same prompt, 400 times, and the outcomes scattered all over the place. Same everything, wildly different runs. The authors had to do a hundred runs per model just to characterize behavior, because a single run tells you almost nothing. Miller's dumb monkey was random on purpose and honest about it. The smart monkey is random and sounds confident every time.

Homogeneous β†’ heterogeneous 🎨

Miller wanted clones. When all you need is volume, identical monkeys are perfect. Two copies of the same model, though? They share the exact same blind spots: they'll stroll right past the same bug sometimes the same way or sometimes in different ways.

XBOW found this out in the most quantified way I've seen. They built "alloy agents": instead of running one model in the loop, they alternate models inside a single conversation thread, a bit of Sonnet, a bit of Gemini, and neither knows the other is there. Success rates went from 25% β†’ 40% β†’ 55%. And the kicker: the more different the two models were (the lower their solve-rate correlation), the bigger the boost. Alloying two models from the same provider did basically nothing: too similar, same blind spots.

That 400-run study backs this from the failure side: the models didn't just differ in how often they won, they failed in completely different ways. One local model kept declaring victory early; one OpenAI model kept burning its whole iteration budget; Gemini was the steady one. Different models, different failure signatures, different blind spots. Which is exactly why you'd want them checking each other. Homogeneity was a feature for monkeys and a bug for agents. Hold that thought.

Honest β†’ lying through its teeth πŸ™ƒ

And then there's the fourth one, which doesn't so much invert as it just… shows up uninvited and ruins the party. And unlike the others, we now have receipts.

FuzzingLabs benchmarked twelve LLMs doing single-pass vulnerability detection on deliberately buggy code. The best model scored under 40% true positives, and every single model threw double-digit false positives. Their own summary: the models are great at spotting risky-looking patterns and terrible at telling which ones are actually real without a human in the loop. That's not a bug detector. That's a very confident intern with a highlighter.

It gets worse when the agent has to judge its own success. In that 400-run study, one model declared the attack COMPLETE before it had actually broken into all the services, in 52 of 100 runs. Over half the time it lied about being done. Not maliciously. It just… had no ground truth telling it otherwise, so it decided it was finished and moved on.

The monkey never lied. The agent lies constantly.

This is the part that took industry embarrassingly long to actually see, and it's the whole ballgame.

That segfault was free ground truth. For the entire memory-corruption bug class, the kernel was the oracle. Nobody had to wonder whether a crash was real, because a crash is machine-checkable, undeniable, and costs nothing to verify. The monkey couldn't fake it if it tried.

An agent's "πŸŽ‰ I FOUND A CRITICAL VULNERABILITY!" with a bunch of πŸŽ‰πŸŽ‰πŸŽ‰ and all caps "I'VE HAD A BREAKTHROUGH" is not that. Not even close. It's not self-verifying at all. And left alone, an agent will happily:

  • edit the source code so its own exploit finally lands, then proudly report the bug it just created,
  • write a test that proves something magnificently tautological, like exec() executes things, therefore: critical RCE,
  • build a beautiful working PoC for a threat model that is complete nonsense ("if an attacker has database write access, they can write to the database"; stunning work, ship it, "GROUND BREAKING"),
  • declare the whole job done when it isn't, per the study above,
  • or, per more than one Devin review, cheerfully push forward on an impossible task rather than admit it's stuck.

Here's the naive loop everybody builds first, with the hole where the kernel used to be:

        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚     task     β”‚
        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
               β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚   LLM agent   β”‚  reads code, forms a theory,
        β”‚  (smart Β· $$) β”‚  adapts... and sometimes fibs
        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
               β–Ό
      "πŸŽ‰ I found a critical RCE!"
               β”‚
               β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚     ????     β”‚   ◄── who checks this?
        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜       the kernel isn't here anymore
               β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚ you, at 2am  β”‚   filtering LIES, not crashes
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

There's a clean way to name the ladder the agent keeps falling off. The ExploitGym benchmark (869 real bugs across userspace programs, V8, and the Linux kernel) frames it right: a vulnerability is not yet an attack. A claim isn't a bug. A bug isn't an exploit. Each rung up that ladder needs proof, and the only honest proof is something that runs. The monkey handed you the top rung for free. The agent hands you a confident sentence and dares you to climb down and check.

The free oracle is gone. Nobody's checking the agent's homework except you.

So the whole job quietly changed category

The other way to look at this:

Babysitting monkeys was a filtering problem. Verification was free; you just filtered the volume. Babysitting agents is a verification problem. The free verification vanished, and now you have to rebuild it yourself.

Once you look at it that way, every serious agent-security system shipping in 2026 suddenly reads the same way, because underneath, they are all doing one thing: manufacturing the oracle that memory-corruption bug hunting always got for free from the kernel.

Look at what the industry is actually building toward. It's uncanny how everyone converges on the same shape:

  • XBOW separates the creative discovery agents from a dedicated validation agent whose entire job is to end-to-end execute the PoC and throw away anything merely theoretical. Discovery proposes; validation decides what's real.
  • Hacktron seems to run on one loud principle: PoC || GTFO. No working proof-of-concept, no finding. That's an oracle: a self-imposed rule that a claim doesn't count until something runs.
  • Cloudflare, in their recent article, goes further and runs discovery and validation on different models, so the thing grading the homework isn't the thing that wrote it. Every finding must ship a PoC that runs against the original, untouched source, so the agent can't cheat by quietly editing the code to make its exploit land. There's your segfault: a crash it can't fake.
  • ExploitGym is what happens when you turn the oracle into a benchmark: give the agent a bug and make it produce a working exploit that achieves real code execution, then check that it actually ran. Pass/fail, machine-checked. The kernel, rebuilt as a scoreboard.
  • The Ralph loop crowd calls this "back-pressure engineering," and Dex Horthy says the quiet part out loud: the best agent engineers spend their days designing the feedback mechanism (deciding exactly how the agent will know whether it succeeded) before they ever hand off the task. That feedback mechanism is the oracle. Everything else is the monkey.

Here's the shape they all land on:

   DISCOVERY                         VERIFICATION
   (creative Β· model A)              (adversarial Β· model B β‰  A)
   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚ hunter agent    β”‚  claim +      β”‚ validator agent        β”‚
   β”‚ must STATE the  │──threat modelβ–Ίβ”‚ β€’ run PoC vs UNTOUCHED  β”‚
   β”‚ threat model    β”‚  + draft PoC  β”‚   source (no edits!)    β”‚
   β”‚ before filing   β”‚               β”‚ β€’ try to DISPROVE it    β”‚
   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β”‚ β€’ never bless your own  β”‚
                                     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                 β”‚ PoC runs?
                                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”
                                        β”‚  yes = the NEW  β”‚
                                        β”‚    segfault     β”‚ ◄── rebuilt oracle
                                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                 β–Ό
                                           β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                           β”‚  human   β”‚
                                           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Notice the threat-model-first bit. That's not bureaucracy. It's how you kill the tautologies before they're born. Make the agent declare who the attacker is or follows the already created threat model and which boundary breaks up front, and "user-with-write-access can write" dies on the spot, because there's no boundary being crossed. You're not filtering that lie after the fact; you're making it un-sayable.

None of this plumbing existed for monkeys. You never needed a second monkey to confirm the first monkey's crash. You absolutely need a second agent to confirm the first agent's claim, because the first agent might simply be making it up.

The army: orchestration, and the babysitting tax nobody warns you about

"Army" was doing a lot of work in Miller's title, and it's doing even more now, because the shape of the army changed too.

Everyone is building an org chart for their agents now, and it is genuinely not settled. The industry is throwing a whole zoo of shapes at the wall, and the differences are worth knowing, because underneath they are all arguments about where the oracle lives:

  • Orchestrator + workers. A lead agent plans, fans out parallel subagents (each with its own context window), then synthesizes, usually with a separate validation pass at the end. Anthropic's research system; XBOW's coordinator spawning isolated "solver" agents, one hunting XSS here, another SQLi there. The oracle is the end-stage validator.
  • Pipeline / assembly line. Fixed stages (recon, hunt, validate, dedup, report) wired as a producer-consumer loop, each stage its own specialized agent. Cloudflare's harness. The oracle is a dedicated station on the line.
  • Generator vs. verifier, on different models. Discovery runs on model A, validation on model B, so nothing grades its own homework. XBOW, Hacktron. This one is the oracle, made structural.
  • Alloy (model-mixing in one thread). Don't add agents, add models: alternate, say, Sonnet and Gemini inside a single loop for complementary blind spots. XBOW. A cheap way to buy heterogeneity.
  • Debate / vote / mixture-of-agents. Ask several models the same thing and have them argue, vote, or defer to a judge. Great for one critical decision, but it multiplies your bill fast, which is why XBOW skipped it in favor of just running more independent agents. Here the oracle is consensus, which (see the next section) is shakier than it looks.
  • Swarm + human review. Many independent agents in isolated sandboxes or git worktrees, each opening a PR for a human to read. Devin, Cursor, the multiplexer crowd, and the far end Steve Yegge and Huntley call "Gas Town," basically Kubernetes for agents. Here you are the oracle, and the bottleneck quietly becomes how many PRs you can review.
  • Single meta-agent in a loop. No committee at all: one decision-maker, fresh context each iteration, sometimes rewriting its own instructions as it goes. Ralph, and the open-source pentest-agent scene (the Cyber-AutoAgent lineage), where more than one builder has reported a single well-run meta-agent beating their multi-agent version. The oracle is whatever feedback the loop is built around.

Two honest notes before you reach for the fanciest option. Anthropic's own writeup says multi-agent shines on broad, parallelizable work and actively hurts on tightly-coupled tasks like coding, where the left hand needs to know what the right hand is doing. And more than one team has found a single sharp agent beats a swarm of mediocre ones. "More agents" is not the axis that matters. Every one of these shapes is just a different answer to the same two questions: where does the oracle live, and who pays the coordination tax.

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  ORCHESTRATOR  β”‚  recon Β· plan Β· spawn
                    β””β”€β”€β”¬β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”¬β”€β”˜
             spawn β–Ό   β–Ό     β–Ό     β–Ό   each isolated Β· own context
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”  Β· different models =
          β”‚ solver A β”‚β”‚solver B β”‚β”‚solver Cβ”‚   different blind spots
          β”‚  SQLi    β”‚β”‚  SSRF  β”‚β”‚  authz β”‚
          β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜β””β”€β”€β”€β”¬β”€β”€β”€β”€β”˜β””β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
               β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
                     β–Ό         β–Ό
               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    the BABYSITTING TAX:
               β”‚  validator (β‰  A)  β”‚    β€’ dedup: 1000s of dupes β†’ 1
               β”‚  PoC || segfault  β”‚    β€’ cost caps: no runaway spawns
               β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β€’ provenance: trace each claim
                         β–Ό                to the agent that made it
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚ human βœ… β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Here's that core loop, since everything below is a variation on it:

 GENERATOR ──► draft ──► VERIFIER ──► accept ──► done
     β–²                       β”‚
     └─── reject + feedback β”€β”€β”˜
 verifier ideally a different model, with explicit criteria

Router (dispatcher). Read the intent, hand off to exactly one specialist:

               β”Œβ”€β”€β–Ί specialist A
 task ──► ROUTER ──► specialist B
               └──► specialist C
 no oracle of its own; it lives in whoever runs

Fan-out / gather. Independent angles at once, merged at the end:

        β”Œβ”€β”€β–Ί worker A ─┐
 task ──┼──► worker B ─┼──► synthesizer ──► out
        └──► worker C β”€β”˜

 the synthesizer is where a verifier bolts on

Agent teams. Like the orchestrator, except the workers don't die after one task; they stay alive and claim work from a shared queue:

 COORDINATOR ──► [ shared queue: β–’ β–’ β–’ ]
                    β–²      β–²      β–²
                   mate   mate   mate    persistent Β· own context
 
 great for parallel long jobs; blind to each other mid-flight

And a few more, drawn small:

Pipeline (assembly line). Fixed stages, each its own agent:

 β”Œβ”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ recon │──►│ hunt  │──►│ validate │──►│ report β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          
                          the oracle is this station

Message bus (event-driven). No fixed order; agents publish and subscribe, a router fans events out to whoever handles them:

 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”  emit  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  route  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ triage  │───────►│  event   │────────►│ net-investigate β”‚
 β”‚ agent   β”‚        β”‚  bus +   β”‚         β”‚ identity        β”‚
 β”‚         │◄───────│  router  │◄────────│ responder       β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  sub   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   sub   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
 
       the oracle has to subscribe like everyone else

Shared state (blackboard). No coordinator at all; agents read and write a common store:

      β”Œβ”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β”
      β”‚agent Aβ”‚      β”‚agent Bβ”‚      β”‚agent Cβ”‚
      β””β”€β”€β”€β”¬β”€β”€β”€β”˜      β””β”€β”€β”€β”¬β”€β”€β”€β”˜      β””β”€β”€β”€β”¬β”€β”€β”€β”˜
          β”‚ r/w          β”‚ r/w          β”‚ r/w
          β–Ό              β–Ό              β–Ό
      β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
      β”‚        shared store (blackboard)      β”‚
      β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      
      
      no referee: needs a "we're done" agent, or it
      loops forever, agents reacting to each other

The shapes differ, but notice the constant: every one has to answer where verification lives. In the pipeline it's a station. On the bus it's a subscriber. In the blackboard it's whichever agent gets to declare "done." You can move the oracle around all you like. You cannot delete it.

But scaling the army adds a whole new babysitting tax that Miller never had to pay, because his monkeys produced honest, self-describing crashes(working PoCs) and these things produce a firehose of overlapping claims:

  • Deduplication becomes its own hard problem. Most of the time, a huge fraction of raw findings are the same underlying bug wearing different hats, thousands of them folding down into far fewer real ones. You can't string-match your way out; deciding whether two logic-flaw writeups are the same bug takes its own agents.
  • Cost can run away silently. Those parallel-subagent systems have no built-in circuit breaker: one agent recursively spawning more agents, or a tool dumping a giant result, can multiply a run's cost by 10x before you notice. Babysitting now includes watching the bill, live.
  • Provenance stops being optional. When one honest monkey crashes, who cares which monkey it was. When one agent out of fifty makes a confident false claim, you need to trace it back. This is the whole reason Cisco's AGNTCY / "Internet of Agents" effort is pouring energy into agent identity: verifiable credentials, an audit trail, the ability to pin every action to the specific agent that took it. In a monkey army that would've been absurd overkill. In a liar army it's table stakes.
  • The loop itself needs a chaperone. That 400-run study's orchestrator is riddled with babysitting logic: detect when the agent is repeating itself, detect when it's stalled in a phase, re-prompt it when it wanders off, kill it when it loops. None of that is the "attack." It's pure supervision, wrapped around a smart thing that keeps getting lost. Miller's monkeys never needed a chaperone inside the loop.

So the org chart is nice, but half the real engineering is the tax underneath it.

Disagreement is the new segfault

Now back to that thought I told you to hold.

In the presentation, Miller wanted a swarm of identical monkeys and just needed volume. The agent army wants the opposite, it wants difference, and once you've internalized the oracle problem, you see why the two facts are actually the same fact.

We lost our free verifier. We have to reconstruct trust from somewhere. And the cheapest honest signal we can manufacture out of a room full of confident, inconsistent liars is where they agree despite being different. One run is a coin flip. But when a model with one set of blind spots and a model with a totally different set both light up on the exact same code, that agreement is expensive (neither could've faked it into the other), and that's your new segfault. Consensus where you had every reason to expect noise. Although, consensus might be caused by convergence, which is total another different issue we have to deal with.

This is why XBOW's alloys work, why Cloudflare cross-checks findings across model families, why the whole field quietly moved from "run the best model" to "run different models and watch the overlap." Is disagreement is the signal now? The correlation of independent, differently-broken judgments is what's left, and you have to engineer the army specifically to produce it. <- seems like a totally wrong take on this convergence.

Homogeneous monkeys, heterogeneous agents. Same word, opposite requirement, and the reason is the missing oracle.

babysitting used to mean counting. now it means refereeing.

So here's where I've landed.

The old job was watching a pile of honest crashes and fishing the good ones out of the noise using tools like Apple's Crashwrangler or other bunch of handcrafted scripts. Tedious, sure, but simple: the signal was trustworthy, there was just a lot of it. Babysitting meant counting.

The new job is refereeing a room full of fast, smart, slightly dishonest toddlers who all swear blind they found something. You make each one prove it against evidence it can't tamper with. You keep discovery and judgment on separate models so nobody grades their own homework. You play them off each other, because the one thing you can trust is where they independently agree. You dedup their overlapping stories, you watch the meter, you keep a receipt for who said what, and you keep a hand on the loop so it doesn't wander into the wall. Babysitting means refereeing.

We babysat the old monkeys because there were millions of them and they were dumb. We babysit the agents because there are a few, they're expensive, and they fib. Same word. Completely different job.

Charlie's monkeys pounded keyboards until something honestly broke. Now they read your source, write a polished report, cite their reasoning, and every so often invent the entire thing.

the half I've been quiet about: knowing where to dig

I've spent this whole post on one of the two hard problems and stayed suspiciously quiet about the other. Everything above is about is it real: the oracle, the verification, the crash that agents can no longer produce for free. But there's a second problem I've barely touched, and it's every bit as hard: where do you even look, classic exploration vs exploitation problem agents face now.

Finding bugs is a search problem. A real target (a codebase, a stripped binary, a baseband) is an enormous space, and almost all of it is boring. The entire game is deciding where to dig, how deep to go, and when to abandon a promising-looking lead and move on. That's the explore-versus-exploit tradeoff, it's old, and dumb fuzzing was genuinely terrible at it: random byte-flipping mostly re-tests the same shallow, easy-to-reach paths over and over. The field's whole answer was coverage-guided fuzzing with GA and evolutionary algorithms: let a coverage signal tell you which mutations actually reached new ground, and steer toward the frontier. (This is exactly why, back before binary code coverage was something you could just switch on, I was bolting Charlie's byte-flipper into my fuzzers, not just for the crashes, but also to get some signal about where in the target I was.) Twenty years of fuzzing progress was mostly this one half.

Agents do not make the search problem easier. They make it bigger. The search space isn't a file format anymore, it's a whole repository plus its dependencies and the entire search space of the model in itself, and the agent has to decide, expensively, which subsystem to read, when to stop enumerating the attack surface and commit to a candidate, when a lead is a dead end. Every one of those calls now costs real money (remember the first inversion). A greedy agent tunnels forever into the first shiny thing it sees; a flighty one skims everything and commits to nothing. Nobody has nailed the balance.

And here I have to be honest, because I've been narrating a solved field and this half isn't one. For the oracle I can point at what the industry builds: PoCs, split validators, disagreement. For search there's no equally clean answer. Coverage-guided fuzzing gave us a gorgeous signal for crashy exploration, but there's nothing nearly as crisp for "am I getting warmer on a logic bug buried in two million lines." This is the part I'm actively messing with right now, and I don't have a tidy answer, just a growing pile of half-working heuristics and a lot of wasted tokens. If you were hoping this post ends with me solving it: sorry. It ends with me admitting it's the interesting problem left.

There's one principle under all of this, and it's worth saying flat out:

A stochastic agent needs a deterministic verifier. The reason the crash worked as an oracle was that it wasn't a matter of opinion. The kernel faulted or it didn't, the same way every time. If you "verify" a random, confident guesser with another random, confident guesser, you haven't rebuilt the oracle; you've just got two dice-rollers who have to agree, and the noise leaks right back in.

anyway

If you're building one of these agent harnesses right now, here's the line I'd stick on the wall:

A smarter monkey was never the hard part. Knowing where it should dig, and whether to believe it when it screams, always was.

Whether to believe it is the oracle: the half that changed, the half that used to be free and now you build. Where it should dig is search: the half that was always hard and just got harder. A smarter generator with no verifier only lies faster; a smarter generator with no sense of where to dig only wanders a bigger maze, at higher cost. You need both, and only one of them was ever free.

So put your money on the two expensive things: knowing what's real, and knowing where to look. Not only on the thing that generates claims. Claims are cheap again, cheaper than they've ever been.

Memory-corruption hunting got its truth for free from the kernel, and coverage-guided fuzzing bought us a clean signal for where to dig. That was the golden age, and it's why the monkeys were beautiful. Neither comes free for the messy logic bugs the agents go hunting now, so we build both. That's the new job, and it's a good one.

Turns out "why is babysitting hard now" is a bigger question than it looks. Go referee your monkeys. And DM me if I got something wrong, or if your agent told you I did. πŸ’


Resources / further reading

  • Charlie Miller, Babysitting an Army of Monkeys (CanSecWest / CodenomiCON, 2010). The OG. Five lines of Python, millions of crashes. Talk on YouTube.
  • Steve Capps' The Monkey (early 80s) and the infinite monkey theorem: where the whole "monkey" thing comes from.
  • XBOW, Agents Built From Alloys. Mixing different-provider models in one thread; the 25β†’40β†’55% result and "more different = better."
  • Hacktron, the PoC || GTFO principle, threat-model-building agents, and their writeups on semantic vs syntactic bugs.
  • Cloudflare, Build your own vulnerability harness. Discovery-vs-validation on different models, PoC-against-untouched-source, dedup-at-scale, and the agent-lying problem, in glorious detail.
  • Anthropic, How we built our multi-agent research system. Orchestrator-worker, per-subagent context windows, and the token-economics reality check (~15x).
  • Geoffrey Huntley / Dex Horthy, the Ralph loop: fresh-context-every-iteration, filesystem-as-memory, and "back-pressure engineering" (a.k.a. designing the oracle before you write the agent).
  • Cisco / AGNTCY, the Internet of Agents: agent identity, provenance, and auditability, the accountability layer for when your army starts lying.
  • Devin / Cursor, parallel autonomous agents in isolated sandboxes/worktrees; the "bottleneck is now how many you can review" problem.
  • FuzzingLabs, Benchmarking LLM agents for vulnerability research. Twelve models, best under 40% true positives, double-digit false positives across the board.
  • ExploitGym (Wang et al., 2026), 869 real vulnerabilities across userspace, V8, and the Linux kernel; can an agent turn a bug into a working exploit? "A vulnerability is not yet an attack."
  • Erdem (2026), How Reliable Are AI Attackers Against a Fixed Vulnerable Target? 400 runs, same target and prompt, wildly different outcomes; premature "COMPLETE" in over half of one model's runs; model-distinctive failure modes.
  • ChiefWiggum, my own Ralph-style Claude Code plugin for iterative vuln hunting: context-build β†’ explore β†’ deep-dive β†’ scrutinize, filesystem memory, REJECT/REQUIRE profiles, team mode. github.com/ant4g0nist/ChiefWiggum.
2026-07-11 Β· ant4g0nist Β· ~26 min Β· 5 tags4 xrefs