Delete the peer count: 7 of 8 colonies recovered instead of 0
An ant colony simulator's ablation table shows the load-balancing mechanism everyone adds to agent pools was the thing making it fragile.
The mechanism that looked essential was the one doing the damage
An ant colony simulator called ant-sim, first written in 2021 and reworked in 2026, was put through a set of ablation runs this year: eight generated worlds, half the foragers killed at minute 50, one mechanism removed at a time. The full rule set recovered from that shock in 0 of 8 worlds. Strip out the crowding estimate — each ant's guess at how many peers are already on a task — and 7 of 8 recovered, in 6.1 minutes. Time spent under-serving foraging fell from 76% to 22%. Tracking error improved too.
The author publishes the commit (d430093), the command that reproduces the table, and the raw outputs, and is explicit about scope: these are simulated ants, not an agent benchmark, and the implications for AI systems are proposals rather than results. An earlier version of the comparison had a confound a reviewer caught, where the "board only" row quietly removed two mechanisms at once. The corrected table separates them.
How the colony actually allocates work
There is no dispatcher. Each ant reads a public board of needs: for every task, how much work has been requested and how much has recently been delivered. It combines that with its own estimate of how crowded each task is, based on the workers it has bumped into, compares its current job with the most pressing alternative, and may switch. Completing work changes the board — food brought home creates a need to store it, brood creates a need for care. The colony's behaviour is the sum of those individual decisions plus the couplings between tasks.
The board is the load-bearing part. Remove it and foraging is under-served 100% of the time in every run, with three of eight worlds starving. Individual response thresholds, the idea that variation between workers damps oscillation, barely moved any metric. The author had a story about why they mattered; the runs did not support it.
Why counting your peers gives you the wrong number
The crowding estimate fails for a reason that transfers directly to software. Ants measure staffing by who they meet, and who you meet depends on where you work. Tasks concentrated in one place looked three to five times more staffed than they were: 4.7x for queen care, 4.4x for digging, 3.9x for storing, 3.6x for brood care. Work spread across the territory looked understaffed — 0.49 for scouting, 0.66 for foraging. The sampling error exists before any ant decides anything.
Moving the count to the nest entrance halved switching and avoided two starvation cases, but barely touched the bias. An entrance samples traffic, not headcount. A nurse working near the shaft crosses it constantly; a forager makes one long trip. Short-trip tasks are over-counted by roughly a factor of four.
Two bugs worth stealing
The board originally divided lifetime need by lifetime delivered work. Early in a run a spike moved the ratio by 17%; three hours later the same spike moved it by 0.003%. Nothing failed visibly — the colony simply went deaf. Decaying both totals with a half-life restored it. The backlog tells you what needs attention; the lifetime count of processed items does not.
The second class of bug: signals generated from a stock rather than from unfinished work. Storage need scaled with everything in the granary, so already-stored food kept demanding storage. Doubling food availability produced smaller colonies — mean population 180 against 206 — as workers piled into storing and digging while brood died. Digging had the same shape: one colony at generation 23 had 140 ants, 93% of its territory excavated, forty mostly empty rooms, 25 diggers and 8 nurses for 58 brood. Tying digging to actual crowding cut nest size to a third and diggers to a quarter with no population loss. The test the author proposes for every signal: what unfinished thing does this measure, and what completed action makes it go away?
One more, on credit. Empty-handed foragers used to score the same as ones carrying food, so the board reported collection was handled while nothing arrived. Making returns count only when food came back left six of eight colonies larger and raised food income in every season. The caveat matters more than the result: a food load is verifiable. For the other tasks, reaching the work site counts as work, and the model cannot tell good output from bad. An agent that writes code needs an independent correctness check; returning is not evidence of success.
Questions You Should Be Asking
- Does your orchestration layer estimate how many agents are on a task, and have you measured that estimate against the true count — or is it sampling traffic, as the entrance-counting version did?
- Which of your routing metrics are lifetime totals? Run the test: how much does a sudden spike move the number at hour one versus hour three?
- For every queue signal you act on: what completed action makes it drop to zero? If none, you are paying agents to redo finished work.
- Can you tell a successful agent run from one that merely returned, or is "came back" your definition of done?
- Has anyone compared this decentralised approach against a plain dispatcher on your workload, with deadlines and the cost of switching agents priced in? The author explicitly has not.
What To Watch Next
The row nobody should ignore is the last one: with neither the crowding estimate nor the sample gate, 73 ants stampeded toward foraging after the shock — four times the herd under the full rule — and all eight colonies recovered in 1.5 minutes, the fastest of any configuration. Herding was fine because it ran in the right direction. Watch for anyone porting this into a real agent pool and reporting what happens when the herd runs the wrong way, and whether a decaying demand board alone beats a dispatcher on cost and latency. Until someone publishes that comparison, this is a well-instrumented hypothesis, not an architecture.
- 1Run one-mechanism-at-a-time ablations under a shock (e.g., kill half your agents at minute 50) — a full rule set that never recovers is a signal, not noise.
- 2Suspect any signal where agents estimate peer counts before acting; test removal, since crowding guesses can lock a system into stale allocations.
- 3Publish the commit hash, the exact reproduce command, and raw outputs alongside the summary table so others can verify your ablation deltas.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
