The Information Machine
Concluded·Day 40·first covered 1 Sep 2026·7 sources·updated 10 Oct 2026

Claude Fable 5.1 on science benchmarks

The gist

Claude Fable 5.1 leads SciUniverse and multiple AI science benchmarks

SciUniverse measures AI performance on physical laboratory execution, not just scientific knowledge, making it a more direct test of whether AI systems can do real research work. Multiple independent benchmarks now place the same model at or near the top on science and autonomous-research tasks, providing cross-validated signal on where frontier capability stands.

The full picture

Claude Fable 5.1 leads the SciUniverse benchmark at a 45.3% pass rate across 92 tasks in chemistry, biology, and materials science, ahead of GPT-5 Astra at 32.5% and Claude Opus 5 at 30.5%; it was also both the highest-performing and most cost-efficient of the three at $40.61 per task. SciUniverse was released by C5R, which built Facility-0, a physical research facility integrating scientific instruments, robotics, and software, in 12 weeks. The benchmark tests whether AI can translate scientific objectives into verifiable lab results, not just recall scientific knowledge. Multiple independent evaluations corroborate Fable 5.1's lead in science-adjacent tasks: Artificial Analysis placed it atop its Intelligence Index at 1,853 Elo with a 62.0% score on SciCode; Vals AI found it first on the RSI Index for autonomous AI research at 35.03%, roughly 21% cheaper than prior leader Claude Opus 5 at $1,481 versus $1,886; and on Terminal-Bench-Science, which tests multi-step scientific workflows including reading papers, running computations, and drawing conclusions, Fable 5.1 scored 52.6%, more than doubling predecessor Fable 5's 24.7%.

How it developed
10 October 2026

C5R released SciUniverse on October 10, a benchmark testing whether AI can turn scientific objectives into verifiable results inside Facility-0, a 12-week-built physical lab, and Claude Fable 5.1 led at 45.3% across 92 tasks at $40.61 per task, ahead of GPT-5 Astra at 32.5% and Opus 5 at 30.5%.

Three evaluations published the same day corroborated the lead: Artificial Analysis's Intelligence Index (1,853 Elo), Vals AI's autonomous-research RSI (35.03%, ~21% cheaper than prior leader), and Terminal-Bench-Science (52.6%, up from Fable 5's 24.7%).

6 October 2026

Google's AIM research agent framework reported, outperforming ScientistOne by up to 4.9 points and reaching equivalent scores up to 3.1x faster.

5 October 2026

Import AI reports SciUniverse results: Claude Fable 5.1 leads at 45.3% pass rate, GPT-5 Astra second, Claude Opus 5 third

24 September 2026

Glitchwire reports C5R built Facility-0 in 12 weeks and released SciUniverse benchmark

1 September 2026

Data Science Dojo reports Fable 5.1 more than doubles Fable 5's Terminal-Bench-Science score (24.7% to 52.6%)

Sources
2 more sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free