Results from tasks both models completed
The tested model IDs are openai/gpt-5.6-luna and deepseek/deepseek-v4-flash. Every result table below comes from a study in which both participated.
A planning score, execution score, defect-recall percentage and simulated strategy return are different measurements. They are not combined into one winner score. Differences in prompt, harness or date can change the result.
How to decide between these versions
Choose the study whose task resembles yours, then read its limitations before using a score gap to choose a model. Compare cost and latency within that same study. A higher mean is not automatically a reliable advantage, and a lower billed cost from a historical run is not a current price quote.
If quality is close or the source reports an inconclusive comparison, test both versions on repeated examples from your own workload. Keep the prompts, tools and grading process the same and inspect invalid outputs as well as successful ones.
Planning models on 29 frozen conversations: What this evaluation measured
GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.
The planner chooses whether to build a plan, dispatch a tool, clarify a request or continue an existing agent. All models are replayed on the same frozen conversation corpus.
Provider-billed per-decision costs and median latency are distinct from quality scores. The highest-scoring planning models were close at the reported sample size.
Planning models on 29 frozen conversations: Published results
| openai/gpt-5.6-luna | 0.731 | 0.068 | $0.0012 | 8.6s |
| deepseek/deepseek-v4-flash | 0.493 | 0.080 | $0.0032 | 12.4s |
No matching rows. Clear the filter to see all records.
16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Published 2026-08-08. Values retain their original units and grading scale.
Planning models on 29 frozen conversations: Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade Router V5 plan replay
Corpus
16-model campaign, 29 frozen conversations; 14 model rows shown in the public table
Metric
Mean graded plan score; standard error reported separately
Planning models on 29 frozen conversations: What the results do not establish
The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.
The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.
The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.
Planning models on 29 frozen conversations: Read the original evidence
Next-action execution model bakeoff: What this evaluation measured
The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.
Models choose the next action inside the same ReAct agent loop. The reported context window in this replay was 73,397 to 83,363 input tokens per decision.
The table reports mean grader score alongside deterministic schema checks, provider billing and production median latency. These columns should be evaluated separately.
Next-action execution model bakeoff: Published results
| openai/gpt-5.6-luna | 89.2 | 99.0% | $1.67 | 5.6s |
| deepseek/deepseek-v4-flash | 75.3 | 98.1% | $6.23 | 17.9s |
No matching rows. Clear the filter to see all records.
23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Published 2026-08-08. Values retain their original units and grading scale.
Next-action execution model bakeoff: Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade Agent V6 next-action replay
Corpus
23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows
Metric
Mean next-action score; schema validity, billed cost and production p50 are separate
Next-action execution model bakeoff: What the results do not establish
Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.
Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.
Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.
Next-action execution model bakeoff: Read the original evidence
Five models completing one sandbox research task: What this evaluation measured
Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.
Each operator had to correlate Strait of Hormuz ship traffic with the next USO session. The model wrote code, ran it, read artifacts and responded to review findings.
The grade trajectory is preserved because a final grade alone hides oscillation and failed remediation. Wall time includes the run workflow rather than only a model response.
Five models completing one sandbox research task: Published results
| openai/gpt-5.6-luna | D → B+ → D → B+ → D → B+ | 59m | $2.63 |
| deepseek/deepseek-v4-flash | D → D → D → D → D → D | 181m | $5.17 |
No matching rows. Clear the filter to see all records.
5 models on the same single research task; real Modal virtual machines. Published 2026-08-08. Values retain their original units and grading scale.
Five models completing one sandbox research task: Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade sandbox operator with the same planner, grader and remediation loop
Corpus
5 models on the same single research task; real Modal virtual machines
Metric
Grade trajectory, elapsed wall time and billed run cost
Five models completing one sandbox research task: What the results do not establish
This is one task, not a broad coding benchmark. A favorable result cannot be generalized to all trading research code.
The early remediation loop oscillated between grades. Later diagnoses found flaws in agent-written audit code; a successful execution status alone did not establish the correctness of delivered analysis.
Flash Lite failed out after 24 minutes and has no reported total cost in the table. Missing cost is not represented as zero.