Evidence for this exact version
The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.
Grok here names the tested underlying model version. Grok Bot is a separate cloud-agent product; a result here does not benchmark its runtime or tools.
Use these results to choose a test
Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.
Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.
Next-action execution model bakeoff
Grok Build 0.1 recorded 84.7 for mean score in the 2026-08-08 publication. 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows.
Study-wide context: The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.
Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.
Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.
Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.
| Mean score | 84.7 |
| Schema valid | 100.0% |
| $ / 1k decisions | $26.88 |
| Prod p50 | 15.4s |
No matching rows. Clear the filter to see all records.