Published model evaluations

GPT 5.6 Luna vs Gemini 3 Flash Preview: Trading Task Evidence

Compare exact model versions on 2 shared published evaluations. See each task's rubric, sample, date and cost units.

As of 2026-08-08

Results from tasks both models completed

The tested model IDs are openai/gpt-5.6-luna and google/gemini-3-flash-preview. Every result table below comes from a study in which both participated.

A planning score, execution score, defect-recall percentage and simulated strategy return are different measurements. They are not combined into one winner score. Differences in prompt, harness or date can change the result.

How to decide between these versions

Choose the study whose task resembles yours, then read its limitations before using a score gap to choose a model. Compare cost and latency within that same study. A higher mean is not automatically a reliable advantage, and a lower billed cost from a historical run is not a current price quote.

If quality is close or the source reports an inconclusive comparison, test both versions on repeated examples from your own workload. Keep the prompts, tools and grading process the same and inspect invalid outputs as well as successful ones.

Next-action execution model bakeoff: What this evaluation measured

The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.

Models choose the next action inside the same ReAct agent loop. The reported context window in this replay was 73,397 to 83,363 input tokens per decision.

The table reports mean grader score alongside deterministic schema checks, provider billing and production median latency. These columns should be evaluated separately.

Next-action execution model bakeoff: Published results

openai/gpt-5.6-luna89.299.0%$1.675.6s
google/gemini-3-flash-preview83.397.1%$24.956.7s

23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Published 2026-08-08. Values retain their original units and grading scale.

Next-action execution model bakeoff: Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade Agent V6 next-action replay

  2. Corpus

    23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows

  3. Metric

    Mean next-action score; schema validity, billed cost and production p50 are separate

Next-action execution model bakeoff: What the results do not establish

Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.

Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.

Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.

Next-action execution model bakeoff: Read the original evidence

Models answering 22 stock questions with SQL: What this evaluation measured

Gemini 3.6 Flash has the highest published average answer score (0.841) and success rate (86.4%) in this six-row SQL comparison. GPT 5.6 Luna scores 0.550 on the same table.

Each question was converted into DuckDB SQL, executed against the real financial lake and graded for answer correctness. This measures executable data retrieval, not free-form ticker opinions.

The source publishes average score, median score, success rate and average execution time. Average execution time is not a production p50 latency measure.

Models answering 22 stock questions with SQL: Published results

google/gemini-3-flash-preview0.7911.0081.8%18.9s
openai/gpt-5.6-luna0.5500.7559.1%27.0s

22 natural-language questions; 6 published model rows. Published 2026-08-08. Values retain their original units and grading scale.

Models answering 22 stock questions with SQL: Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade financial lake SQL screener, descended from EvaluateGPT

  2. Corpus

    22 natural-language questions; 6 published model rows

  3. Metric

    Average answer score, success rate and average execution time

Models answering 22 stock questions with SQL: What the results do not establish

These are the August study’s exact versions, data and questions. They do not establish the best model for every stock-research task or current production defaults.

A successful SQL query is not a profitable investment strategy. No return or trading signal performance is measured in this table.

Answer scores depend on the question corpus and grading process. Changing schemas, tools or prompts can change the ranking.

Models answering 22 stock questions with SQL: Read the original evidence