Intology published results on August 13 showing its Locus agent, running on Opus 5, scored 51.6% on PostTrainBench+, clearing the 51.1% human baseline at a cost of more than 4,000 H100-hours; on the standard benchmark, Locus scored 44.7% against 34.1% for Opus 5 without the harness.
A second source corroborated the result, added competitor scores (Opus 4.8 at 44.3% and GLM 5.2 at 42.7% on PostTrainBench+), and characterized the compute cost as proof-of-concept territory rather than a practical research workflow.