Published model evaluations

Grok 4.20: Trading Task Evaluations

See how x-ai/grok-4.20 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-04-06

Exact tested ID
x-ai/grok-4.20
Published studies
1
Provider
x-ai

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

Grok here names the tested underlying model version. Grok Bot is a separate cloud-agent product; a result here does not benchmark its runtime or tools.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Eleven models building an options strategy

Grok 4.20 recorded 4 for score in the 2026-04-06 publication. 11 model runs; same task, account context, tools and 25-iteration budget.

Study-wide context: Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective.

The publication presents one run per model in this comparison. It does not establish seed-to-seed uncertainty or a general model-quality ranking.

The scorecard and narrative disagree on some subagent/portfolio counts and on how consistently the selected Flash strategy was positive across regimes. The table preserves the published scorecard; conflicting process counts and regime claims are not resolved by inference.

The article describes historical simulations and deployment actions. This study page does not assert independently verified live account returns, achieved doubling, or a deployment recommendation.

Score4
Verdictfail
Subagents2
Portfolios11
Time10.3m

Continue exploring