Published model evaluations

GPT 5.6 Luna: Trading Task Evaluations

See how openai/gpt-5.6-luna performed in 5 trading task evaluations, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
openai/gpt-5.6-luna
Published studies
5
Provider
openai

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

The provider name identifies the model family. The tools and harness around the model determine what it can do.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Planning models on 29 frozen conversations

GPT 5.6 Luna recorded 0.731 for score in the 2026-08-08 publication. 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table.

Study-wide context: GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.

The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.

The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.

The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.

Score0.731
+/- SE0.068
$ / decision$0.0012
p508.6s

Next-action execution model bakeoff

GPT 5.6 Luna recorded 89.2 for mean score in the 2026-08-08 publication. 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows.

Study-wide context: The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.

Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.

Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.

Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.

Mean score89.2
Schema valid99.0%
$ / 1k decisions$1.67
Prod p505.6s

Artifact grading and repeatability study

GPT 5.6 Luna recorded 83.1% for defects found in the 2026-08-08 publication. 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table.

Study-wide context: Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.

A reducer cannot discover a defect absent from its mapper inputs. These results do not establish end-to-end recall for every arbitrary artifact.

GPT 5.6 Luna and Tencent Hy3 tie on reported defect recall (83.1%). Their repeated-grade counts differ by one in 23 tests; that is not evidence of a universal review advantage.

This is a task-specific snapshot. Current grader defaults and newer model versions are not inferred from an August publication.

Defects found83.1%
Same grade on re-run19 of 23

Five models completing one sandbox research task

GPT 5.6 Luna recorded D → B+ → D → B+ → D → B+ for grade trajectory in the 2026-08-08 publication. 5 models on the same single research task; real Modal virtual machines.

Study-wide context: Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.

This is one task, not a broad coding benchmark. A favorable result cannot be generalized to all trading research code.

The early remediation loop oscillated between grades. Later diagnoses found flaws in agent-written audit code; a successful execution status alone did not establish the correctness of delivered analysis.

Flash Lite failed out after 24 minutes and has no reported total cost in the table. Missing cost is not represented as zero.

Grade trajectoryD → B+ → D → B+ → D → B+
Wall59m
Cost$2.63

Models answering 22 stock questions with SQL

GPT 5.6 Luna recorded 0.550 for average score in the 2026-08-08 publication. 22 natural-language questions; 6 published model rows.

Study-wide context: Gemini 3.6 Flash has the highest published average answer score (0.841) and success rate (86.4%) in this six-row SQL comparison. GPT 5.6 Luna scores 0.550 on the same table.

These are the August study’s exact versions, data and questions. They do not establish the best model for every stock-research task or current production defaults.

A successful SQL query is not a profitable investment strategy. No return or trading signal performance is measured in this table.

Answer scores depend on the question corpus and grading process. Changing schemas, tools or prompts can change the ranking.

Average score0.550
Median0.75
Success rate59.1%
Avg execution27.0s

Continue exploring