Public page collection

Benchmark studies: published pages

Browse the published Benchmark studies collection with direct links and numbered pages.

Published pages
6
Directory page
1 of 1

Pages 1–6

Artifact grading and repeatability study

Published 2026-08-08: 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table. Defect recall and repeated-grade agreement. See the results, how they were measured and where the findings stop.

Open page

Next-action execution model bakeoff

Published 2026-08-08: 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Mean next-action score; schema validity, billed cost and production p50 are separate. See the results, how they were measured and where the findings stop.

Open page

Eleven models building an options strategy

Published 2026-04-06: 11 model runs; same task, account context, tools and 25-iteration budget. Goal-oriented evaluator score / 100. See the results, how they were measured and where the findings stop.

Open page

Planning models on 29 frozen conversations

Published 2026-08-08: 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Mean graded plan score; standard error reported separately. See the results, how they were measured and where the findings stop.

Open page

Five models completing one sandbox research task

Published 2026-08-08: 5 models on the same single research task; real Modal virtual machines. Grade trajectory, elapsed wall time and billed run cost. See the results, how they were measured and where the findings stop.

Open page

Models answering 22 stock questions with SQL

Published 2026-08-08: 22 natural-language questions; 6 published model rows. Average answer score, success rate and average execution time. See the results, how they were measured and where the findings stop.

Open page

Continue browsing