Published model evaluations

Best AI Models for Options Strategy Development: Published Evaluations

Which models performed well on options strategy development in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

As of 2026-04-06

What the options strategy development evidence supports

As published on 2026-04-06: Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective.

Test against your own workload

Check whether the tested task and required output match your workload. Use a fixed set of representative examples, repeat the runs and inspect failures. Compare cost and latency alongside quality.

These are published historical results. Internal customer conversations remain private. The results do not establish the current production default or promise profitable trading.

What this evaluation measured

Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective.

Each run received the same $25,000 account objective, watchlist context, available tools and iteration budget. The run had to research, generate configurations, backtest across regimes and make a deployment decision.

The published rubric combines strategy fitness (40%), evidence strength (25%), exploration coverage (20%) and risk realism (15%). A high score reflects this rubric and objective, not a realized investment return.

Published results

Gemini 3 Flash Preview66mixed3645.5m
Gemini 3.1 Flash Lite53weak065.3m
Claude Sonnet 4.650weak2532m
Gemini 3 Pro Preview50weak2428m
Gemma 4 31B50weak03~40m
MiMo v2 Pro47weak0418m
Kimi K2.55fail0820.4m
Grok 4.204fail21110.3m
GPT-5.4-mini1fail005.3m
GPT-5.40fail0010.7m
Mistral Small 26030fail005.4m

11 model runs; same task, account context, tools and 25-iteration budget. Published 2026-04-06. Values retain their original units and grading scale.

Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade Aurora autonomous research and options workflow; Agent Run Evaluator V6

  2. Corpus

    11 model runs; same task, account context, tools and 25-iteration budget

  3. Metric

    Goal-oriented evaluator score / 100

What the results do not establish

The publication presents one run per model in this comparison. It does not establish seed-to-seed uncertainty or a general model-quality ranking.

The scorecard and narrative disagree on some subagent/portfolio counts and on how consistently the selected Flash strategy was positive across regimes. The table preserves the published scorecard; conflicting process counts and regime claims are not resolved by inference.

The article describes historical simulations and deployment actions. This study page does not assert independently verified live account returns, achieved doubling, or a deployment recommendation.

Read the original evidence

Continue exploring