What the agent planning evidence supports
As published on 2026-08-08: GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.
Test against your own workload
Check whether the tested task and required output match your workload. Use a fixed set of representative examples, repeat the runs and inspect failures. Compare cost and latency alongside quality.
These are published historical results. Internal customer conversations remain private. The results do not establish the current production default or promise profitable trading.
What this evaluation measured
GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.
The planner chooses whether to build a plan, dispatch a tool, clarify a request or continue an existing agent. All models are replayed on the same frozen conversation corpus.
Provider-billed per-decision costs and median latency are distinct from quality scores. The highest-scoring planning models were close at the reported sample size.
Published results
| meta/muse-spark-1.1 | 0.759 | 0.079 | $0.0229 | 9.6s |
| openai/gpt-5.6-luna | 0.731 | 0.068 | $0.0012 | 8.6s |
| openai/gpt-5.6-luna-pro | 0.731 | 0.068 | $0.0076 | 21.9s |
| moonshotai/kimi-k3 | 0.731 | 0.068 | $0.0799 | 38.9s |
| x-ai/grok-4.5 | 0.724 | 0.065 | $0.0519 | 23.5s |
| deepseek/deepseek-v4-pro | 0.710 | 0.064 | $0.0142 | 27.2s |
| google/gemini-3.6-flash | 0.662 | 0.070 | $0.0387 | 8.5s |
| anthropic/claude-opus-5 | 0.662 | 0.074 | $0.3159 | 26.6s |
| openai/gpt-5.6-terra-pro | 0.634 | 0.073 | $0.0646 | 16.0s |
| google/gemini-3.1-flash-lite | 0.586 | 0.066 | $0.0031 | 1.8s |
| anthropic/claude-sonnet-5 | 0.579 | 0.071 | $0.1348 | 24.9s |
| z-ai/glm-5.2 | 0.515 | 0.079 | $0.0246 | 44.8s |
| deepseek/deepseek-v4-flash | 0.493 | 0.080 | $0.0032 | 12.4s |
| qwen/qwen3.8-max | 0.400 | 0.082 | $0.0421 | 168.7s |
No matching rows. Clear the filter to see all records.
16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Published 2026-08-08. Values retain their original units and grading scale.
Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade Router V5 plan replay
Corpus
16-model campaign, 29 frozen conversations; 14 model rows shown in the public table
Metric
Mean graded plan score; standard error reported separately
What the results do not establish
The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.
The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.
The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.