Published model evaluations

GPT 5.6 Luna vs Claude Opus 5: Trading Task Evidence

Compare exact model versions on 1 shared published evaluation. See each task's rubric, sample, date and cost units.

As of 2026-08-08

Results from tasks both models completed

The tested model IDs are openai/gpt-5.6-luna and anthropic/claude-opus-5. Every result table below comes from a study in which both participated.

A planning score, execution score, defect-recall percentage and simulated strategy return are different measurements. They are not combined into one winner score. Differences in prompt, harness or date can change the result.

How to decide between these versions

Choose the study whose task resembles yours, then read its limitations before using a score gap to choose a model. Compare cost and latency within that same study. A higher mean is not automatically a reliable advantage, and a lower billed cost from a historical run is not a current price quote.

If quality is close or the source reports an inconclusive comparison, test both versions on repeated examples from your own workload. Keep the prompts, tools and grading process the same and inspect invalid outputs as well as successful ones.

Planning models on 29 frozen conversations: What this evaluation measured

GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.

The planner chooses whether to build a plan, dispatch a tool, clarify a request or continue an existing agent. All models are replayed on the same frozen conversation corpus.

Provider-billed per-decision costs and median latency are distinct from quality scores. The highest-scoring planning models were close at the reported sample size.

Planning models on 29 frozen conversations: Published results

openai/gpt-5.6-luna0.7310.068$0.00128.6s
anthropic/claude-opus-50.6620.074$0.315926.6s

16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Published 2026-08-08. Values retain their original units and grading scale.

Planning models on 29 frozen conversations: Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade Router V5 plan replay

  2. Corpus

    16-model campaign, 29 frozen conversations; 14 model rows shown in the public table

  3. Metric

    Mean graded plan score; standard error reported separately

Planning models on 29 frozen conversations: What the results do not establish

The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.

The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.

The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.

Planning models on 29 frozen conversations: Read the original evidence