What the sandbox code generation evidence supports
As published on 2026-08-08: Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.
Test against your own workload
Check whether the tested task and required output match your workload. Use a fixed set of representative examples, repeat the runs and inspect failures. Compare cost and latency alongside quality.
These are published historical results. Internal customer conversations remain private. The results do not establish the current production default or promise profitable trading.
What this evaluation measured
Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.
Each operator had to correlate Strait of Hormuz ship traffic with the next USO session. The model wrote code, ran it, read artifacts and responded to review findings.
The grade trajectory is preserved because a final grade alone hides oscillation and failed remediation. Wall time includes the run workflow rather than only a model response.
Published results
| openai/gpt-5.6-luna | D → B+ → D → B+ → D → B+ | 59m | $2.63 |
| z-ai/glm-5.2 | D → D → B+ → D → D → D | 91m | $2.72 |
| qwen/qwen3.7-flash | D → D → D | 124m | $2.85 |
| deepseek/deepseek-v4-flash | D → D → D → D → D → D | 181m | $5.17 |
| google/gemini-3.5-flash-lite | D, then failed out | 24m | — |
No matching rows. Clear the filter to see all records.
5 models on the same single research task; real Modal virtual machines. Published 2026-08-08. Values retain their original units and grading scale.
Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade sandbox operator with the same planner, grader and remediation loop
Corpus
5 models on the same single research task; real Modal virtual machines
Metric
Grade trajectory, elapsed wall time and billed run cost
What the results do not establish
This is one task, not a broad coding benchmark. A favorable result cannot be generalized to all trading research code.
The early remediation loop oscillated between grades. Later diagnoses found flaws in agent-written audit code; a successful execution status alone did not establish the correctness of delivered analysis.
Flash Lite failed out after 24 minutes and has no reported total cost in the table. Missing cost is not represented as zero.