Results from tasks both models completed
The tested model IDs are google/gemini-3-flash-preview and x-ai/grok-4.20. Every result table below comes from a study in which both participated.
A planning score, execution score, defect-recall percentage and simulated strategy return are different measurements. They are not combined into one winner score. Differences in prompt, harness or date can change the result.
How to decide between these versions
Choose the study whose task resembles yours, then read its limitations before using a score gap to choose a model. Compare cost and latency within that same study. A higher mean is not automatically a reliable advantage, and a lower billed cost from a historical run is not a current price quote.
If quality is close or the source reports an inconclusive comparison, test both versions on repeated examples from your own workload. Keep the prompts, tools and grading process the same and inspect invalid outputs as well as successful ones.
Eleven models building an options strategy: What this evaluation measured
Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective.
Each run received the same $25,000 account objective, watchlist context, available tools and iteration budget. The run had to research, generate configurations, backtest across regimes and make a deployment decision.
The published rubric combines strategy fitness (40%), evidence strength (25%), exploration coverage (20%) and risk realism (15%). A high score reflects this rubric and objective, not a realized investment return.
Eleven models building an options strategy: Published results
| Gemini 3 Flash Preview | 66 | mixed | 3 | 6 | 45.5m |
| Grok 4.20 | 4 | fail | 2 | 11 | 10.3m |
No matching rows. Clear the filter to see all records.
11 model runs; same task, account context, tools and 25-iteration budget. Published 2026-04-06. Values retain their original units and grading scale.
Eleven models building an options strategy: Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade Aurora autonomous research and options workflow; Agent Run Evaluator V6
Corpus
11 model runs; same task, account context, tools and 25-iteration budget
Metric
Goal-oriented evaluator score / 100
Eleven models building an options strategy: What the results do not establish
The publication presents one run per model in this comparison. It does not establish seed-to-seed uncertainty or a general model-quality ranking.
The scorecard and narrative disagree on some subagent/portfolio counts and on how consistently the selected Flash strategy was positive across regimes. The table preserves the published scorecard; conflicting process counts and regime claims are not resolved by inference.
The article describes historical simulations and deployment actions. This study page does not assert independently verified live account returns, achieved doubling, or a deployment recommendation.