Published model evaluations

Best AI Models for Sandbox Code Generation: Published Evaluations

Which models performed well on sandbox code generation in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

As of 2026-08-08

What the sandbox code generation evidence supports

As published on 2026-08-08: Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.

Test against your own workload

Check whether the tested task and required output match your workload. Use a fixed set of representative examples, repeat the runs and inspect failures. Compare cost and latency alongside quality.

These are published historical results. Internal customer conversations remain private. The results do not establish the current production default or promise profitable trading.

What this evaluation measured

Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.

Each operator had to correlate Strait of Hormuz ship traffic with the next USO session. The model wrote code, ran it, read artifacts and responded to review findings.

The grade trajectory is preserved because a final grade alone hides oscillation and failed remediation. Wall time includes the run workflow rather than only a model response.

Published results

openai/gpt-5.6-lunaD → B+ → D → B+ → D → B+59m$2.63
z-ai/glm-5.2D → D → B+ → D → D → D91m$2.72
qwen/qwen3.7-flashD → D → D124m$2.85
deepseek/deepseek-v4-flashD → D → D → D → D → D181m$5.17
google/gemini-3.5-flash-liteD, then failed out24m—

5 models on the same single research task; real Modal virtual machines. Published 2026-08-08. Values retain their original units and grading scale.

Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade sandbox operator with the same planner, grader and remediation loop

  2. Corpus

    5 models on the same single research task; real Modal virtual machines

  3. Metric

    Grade trajectory, elapsed wall time and billed run cost

What the results do not establish

This is one task, not a broad coding benchmark. A favorable result cannot be generalized to all trading research code.

The early remediation loop oscillated between grades. Later diagnoses found flaws in agent-written audit code; a successful execution status alone did not establish the correctness of delivered analysis.

Flash Lite failed out after 24 minutes and has no reported total cost in the table. Missing cost is not represented as zero.

Read the original evidence

Continue exploring