Published model evaluations

Best AI Models for Algorithmic Trading: Task-by-Task Evidence

Compare published model evaluations for options strategies, planning, agent execution, artifact review, sandbox code and stock SQL screening. Each task keeps its own methodology.

As of 2026-08-08

Published task studies
6
Latest publication
2026-08-08
Live-return leaderboard
Not measured

Choose the model for the job

The published evidence does not identify one universal best trading model. The options-workflow study favored a Google Flash model on its rubric; the later planning and execution studies favored a cost-efficient OpenAI version for different reasons; SQL screening favored a different Google version.

A model supplies reasoning inside a harness. Data access, executable strategy rules, simulation infrastructure, schema validation, iteration limits and review determine whether it can actually complete a trading task. Meta Muse and Grok Bot are agent products; Muse Spark and Grok model versions are evaluated separately here.

Published findings by task

Options Strategy Development2026-04-06Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective.
Agent Planning2026-08-08GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.
Agent Execution2026-08-08The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.
Artifact Review2026-08-08Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.
Sandbox Code Generation2026-08-08Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.
Stock Screening2026-08-08Gemini 3.6 Flash has the highest published average answer score (0.841) and success rate (86.4%) in this six-row SQL comparison. GPT 5.6 Luna scores 0.550 on the same table.

How to compare the evidence

  1. Same task and harness

    Compare models on the same inputs, tools and output contract. Record the exact version IDs.

  2. More than one favorable result

    Report repeated samples, uncertainty and failure behavior when available. A one-task run is a screen, not a universal verdict.

  3. Separate measurements

    Quality, schema validity, elapsed time, provider-billed cost and trading returns have different units.

  4. Inspect the judge

    The published execution regrade showed vendor-dependent score gaps. Calibration and deterministic checks matter.

Historical evidence and current decisions

The results below retain their publication dates and original units. August deployment decisions are historical observations, not a claim about current defaults. Prices and model availability may have changed.

Evaluator confidence, historical backtest returns, a deployment action and verified live account performance measure different things. Evidence for one does not establish the others.