Published model evaluations

Five models completing one sandbox research task

Published 2026-08-08: 5 models on the same single research task; real Modal virtual machines. Grade trajectory, elapsed wall time and billed run cost. See the results, how they were measured and where the findings stop.

As of 2026-08-08

Published
2026-08-08
Published model rows
5
Evidence
Historical task evaluation

What this evaluation measured

Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.

Each operator had to correlate Strait of Hormuz ship traffic with the next USO session. The model wrote code, ran it, read artifacts and responded to review findings.

The grade trajectory is preserved because a final grade alone hides oscillation and failed remediation. Wall time includes the run workflow rather than only a model response.

Published results

openai/gpt-5.6-lunaD → B+ → D → B+ → D → B+59m$2.63
z-ai/glm-5.2D → D → B+ → D → D → D91m$2.72
qwen/qwen3.7-flashD → D → D124m$2.85
deepseek/deepseek-v4-flashD → D → D → D → D → D181m$5.17
google/gemini-3.5-flash-liteD, then failed out24m—

5 models on the same single research task; real Modal virtual machines. Published 2026-08-08. Values retain their original units and grading scale.

Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade sandbox operator with the same planner, grader and remediation loop

  2. Corpus

    5 models on the same single research task; real Modal virtual machines

  3. Metric

    Grade trajectory, elapsed wall time and billed run cost

What the results do not establish

This is one task, not a broad coding benchmark. A favorable result cannot be generalized to all trading research code.

The early remediation loop oscillated between grades. Later diagnoses found flaws in agent-written audit code; a successful execution status alone did not establish the correctness of delivered analysis.

Flash Lite failed out after 24 minutes and has no reported total cost in the table. Missing cost is not represented as zero.

Read the original evidence

Continue exploring