Choose the model for the job
The published evidence does not identify one universal best trading model. The options-workflow study favored a Google Flash model on its rubric; the later planning and execution studies favored a cost-efficient OpenAI version for different reasons; SQL screening favored a different Google version.
A model supplies reasoning inside a harness. Data access, executable strategy rules, simulation infrastructure, schema validation, iteration limits and review determine whether it can actually complete a trading task. Meta Muse and Grok Bot are agent products; Muse Spark and Grok model versions are evaluated separately here.
Published findings by task
| Options Strategy Development | 2026-04-06 | Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective. |
| Agent Planning | 2026-08-08 | GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims. |
| Agent Execution | 2026-08-08 | The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost. |
| Artifact Review | 2026-08-08 | Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23. |
| Sandbox Code Generation | 2026-08-08 | Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations. |
| Stock Screening | 2026-08-08 | Gemini 3.6 Flash has the highest published average answer score (0.841) and success rate (86.4%) in this six-row SQL comparison. GPT 5.6 Luna scores 0.550 on the same table. |
No matching rows. Clear the filter to see all records.
How to compare the evidence
Same task and harness
Compare models on the same inputs, tools and output contract. Record the exact version IDs.
More than one favorable result
Report repeated samples, uncertainty and failure behavior when available. A one-task run is a screen, not a universal verdict.
Separate measurements
Quality, schema validity, elapsed time, provider-billed cost and trading returns have different units.
Inspect the judge
The published execution regrade showed vendor-dependent score gaps. Calibration and deterministic checks matter.
Historical evidence and current decisions
The results below retain their publication dates and original units. August deployment decisions are historical observations, not a claim about current defaults. Prices and model availability may have changed.
Evaluator confidence, historical backtest returns, a deployment action and verified live account performance measure different things. Evidence for one does not establish the others.