Published model evaluations

Grok Build 0.1: Trading Task Evaluations

See how x-ai/grok-build-0.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
x-ai/grok-build-0.1
Published studies
1
Provider
x-ai

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

Grok here names the tested underlying model version. Grok Bot is a separate cloud-agent product; a result here does not benchmark its runtime or tools.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Next-action execution model bakeoff

Grok Build 0.1 recorded 84.7 for mean score in the 2026-08-08 publication. 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows.

Study-wide context: The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.

Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.

Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.

Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.

Mean score84.7
Schema valid100.0%
$ / 1k decisions$26.88
Prod p5015.4s

Continue exploring