Pages 1–6
Artifact grading and repeatability study
Published 2026-08-08: 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table. Defect recall and repeated-grade agreement. See the results, how they were measured and where the findings stop.
Open pageNext-action execution model bakeoff
Published 2026-08-08: 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Mean next-action score; schema validity, billed cost and production p50 are separate. See the results, how they were measured and where the findings stop.
Open pageEleven models building an options strategy
Published 2026-04-06: 11 model runs; same task, account context, tools and 25-iteration budget. Goal-oriented evaluator score / 100. See the results, how they were measured and where the findings stop.
Open pagePlanning models on 29 frozen conversations
Published 2026-08-08: 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Mean graded plan score; standard error reported separately. See the results, how they were measured and where the findings stop.
Open pageFive models completing one sandbox research task
Published 2026-08-08: 5 models on the same single research task; real Modal virtual machines. Grade trajectory, elapsed wall time and billed run cost. See the results, how they were measured and where the findings stop.
Open pageModels answering 22 stock questions with SQL
Published 2026-08-08: 22 natural-language questions; 6 published model rows. Average answer score, success rate and average execution time. See the results, how they were measured and where the findings stop.
Open page