C5R released SciUniverse on October 10, a benchmark testing whether AI can turn scientific objectives into verifiable results inside Facility-0, a 12-week-built physical lab, and Claude Fable 5.1 led at 45.3% across 92 tasks at $40.61 per task, ahead of GPT-5 Astra at 32.5% and Opus 5 at 30.5%.
Three evaluations published the same day corroborated the lead: Artificial Analysis's Intelligence Index (1,853 Elo), Vals AI's autonomous-research RSI (35.03%, ~21% cheaper than prior leader), and Terminal-Bench-Science (52.6%, up from Fable 5's 24.7%).