A Shelf Life Shorter Than Peer Review
AgentRadio makes an unusually clean argument. Take a multi-agent setup where agents wait to be mentioned by each other, and change exactly one thing: run that wait as a background task rather than a foreground one, so an agent notices a colleague's message while it is working instead of only when it stops to look.
That single swap is worth +10.5 points on one model and +11.3 on another, with proper significance testing behind it. One variable, isolated, measured. I wish more papers did this.
The closing claim is what got attention. Four coordinated agents on the older model score 62.1%; a single agent on the newer one scores 57.2%. The authors conclude they have found a degree of freedom worth more than a model generation.
Nine days later I checked the same public leaderboard. A single agent on the next model scores 63.17%.
One agent had caught up with four, at a fraction of the cost, before the preprint was a fortnight old.
Be careful with that, though, because it is exactly the kind of line that is too satisfying to check. The interval on that score is ±5, which makes the two results statistically indistinguishable, and nobody has run AgentRadio on the newer model — it would presumably gain too. This is not a refutation.
It is a shelf-life warning. Any claim shaped like our scaffolding beats a model generation is measured against a baseline that moves, and this one moved inside two weeks. The finding about background waiting is durable. The headline comparison was decaying before it was published.
There is a cost most coverage skipped, too. Passive awareness runs about 25% more expensive per task, and against 47 improved results there were 23 regressions — agents pulled off productive work by a colleague's message arriving mid-thought. The paper has no limitations section in which to say so.
I want the mechanism. I do not want the comparison.