What this evaluation measured
Gemini 3.6 Flash has the highest published average answer score (0.841) and success rate (86.4%) in this six-row SQL comparison. GPT 5.6 Luna scores 0.550 on the same table.
Each question was converted into DuckDB SQL, executed against the real financial lake and graded for answer correctness. This measures executable data retrieval, not free-form ticker opinions.
The source publishes average score, median score, success rate and average execution time. Average execution time is not a production p50 latency measure.
Published results
| google/gemini-3.6-flash | 0.841 | 1.00 | 86.4% | 26.6s |
| google/gemini-3-flash-preview | 0.791 | 1.00 | 81.8% | 18.9s |
| google/gemini-3.5-flash | 0.727 | 1.00 | 72.7% | 33.3s |
| google/gemini-3.5-flash-lite | 0.664 | 0.90 | 72.7% | 19.4s |
| openai/gpt-5-mini | 0.664 | 1.00 | 68.2% | 41.5s |
| openai/gpt-5.6-luna | 0.550 | 0.75 | 59.1% | 27.0s |
No matching rows. Clear the filter to see all records.
22 natural-language questions; 6 published model rows. Published 2026-08-08. Values retain their original units and grading scale.
Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
NexusTrade financial lake SQL screener, descended from EvaluateGPT
Corpus
22 natural-language questions; 6 published model rows
Metric
Average answer score, success rate and average execution time
What the results do not establish
These are the August study’s exact versions, data and questions. They do not establish the best model for every stock-research task or current production defaults.
A successful SQL query is not a profitable investment strategy. No return or trading signal performance is measured in this table.
Answer scores depend on the question corpus and grading process. Changing schemas, tools or prompts can change the ranking.