Evidence for this exact version
The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.
The provider name identifies the model family. The tools and harness around the model determine what it can do.
Use these results to choose a test
Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.
Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.
Eleven models building an options strategy
Gemini 3 Flash Preview recorded 66 for score in the 2026-04-06 publication. 11 model runs; same task, account context, tools and 25-iteration budget.
Study-wide context: Gemini 3 Flash Preview has the highest published score, 66/100. The evaluator still calls that mixed and asks for more iteration; no result proves the account-doubling objective.
The publication presents one run per model in this comparison. It does not establish seed-to-seed uncertainty or a general model-quality ranking.
The scorecard and narrative disagree on some subagent/portfolio counts and on how consistently the selected Flash strategy was positive across regimes. The table preserves the published scorecard; conflicting process counts and regime claims are not resolved by inference.
The article describes historical simulations and deployment actions. This study page does not assert independently verified live account returns, achieved doubling, or a deployment recommendation.
| Score | 66 |
| Verdict | mixed |
| Subagents | 3 |
| Portfolios | 6 |
| Time | 45.5m |
No matching rows. Clear the filter to see all records.
Next-action execution model bakeoff
Gemini 3 Flash Preview recorded 83.3 for mean score in the 2026-08-08 publication. 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows.
Study-wide context: The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.
Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.
Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.
Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.
| Mean score | 83.3 |
| Schema valid | 97.1% |
| $ / 1k decisions | $24.95 |
| Prod p50 | 6.7s |
No matching rows. Clear the filter to see all records.
Models answering 22 stock questions with SQL
Gemini 3 Flash Preview recorded 0.791 for average score in the 2026-08-08 publication. 22 natural-language questions; 6 published model rows.
Study-wide context: Gemini 3.6 Flash has the highest published average answer score (0.841) and success rate (86.4%) in this six-row SQL comparison. GPT 5.6 Luna scores 0.550 on the same table.
These are the August study’s exact versions, data and questions. They do not establish the best model for every stock-research task or current production defaults.
A successful SQL query is not a profitable investment strategy. No return or trading signal performance is measured in this table.
Answer scores depend on the question corpus and grading process. Changing schemas, tools or prompts can change the ranking.
| Average score | 0.791 |
| Median | 1.00 |
| Success rate | 81.8% |
| Avg execution | 18.9s |
No matching rows. Clear the filter to see all records.