Four papers published October 7 examine when multi-agent AI helps or hurts depending on task structure.
Meta's RankEvolve found cross-family code review raised correct patches from 45.8% to 62.5%. A Stanford paper found a single coordinating agent captured 64% of achievable value on contested resources versus 30% for Opus 5 multi-agent teams. Microsoft's Agensh showed leaderless coding teams scale to 128 agents on ProgramBench, while Toby Ord found 10x more agents yields only 3x to 5x gains.