Evidence for this exact version
The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.
The provider name identifies the model family. The tools and harness around the model determine what it can do.
Use these results to choose a test
Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.
Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.
Planning models on 29 frozen conversations
Deepseek V4 Pro recorded 0.710 for score in the 2026-08-08 publication. 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table.
Study-wide context: GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.
The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.
The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.
The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.
| Score | 0.710 |
| +/- SE | 0.064 |
| $ / decision | $0.0142 |
| p50 | 27.2s |
No matching rows. Clear the filter to see all records.