Published model evaluations

Grok 4.5: Trading Task Evaluations

See how x-ai/grok-4.5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
x-ai/grok-4.5
Published studies
1
Provider
x-ai

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

Grok here names the tested underlying model version. Grok Bot is a separate cloud-agent product; a result here does not benchmark its runtime or tools.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Planning models on 29 frozen conversations

Grok 4.5 recorded 0.724 for score in the 2026-08-08 publication. 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table.

Study-wide context: GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.

The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.

The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.

The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.

Score0.724
+/- SE0.065
$ / decision$0.0519
p5023.5s

Continue exploring