Results from tasks both models completed
The tested model IDs are openai/gpt-5.6-luna and minimax/minimax-m3. Every result table below comes from a study in which both participated.
A planning score, execution score, defect-recall percentage and simulated strategy return are different measurements. They are not combined into one winner score. Differences in prompt, harness or date can change the result.
How to decide between these versions
Choose the study whose task resembles yours, then read its limitations before using a score gap to choose a model. Compare cost and latency within that same study. A higher mean is not automatically a reliable advantage, and a lower billed cost from a historical run is not a current price quote.
If quality is close or the source reports an inconclusive comparison, test both versions on repeated examples from your own workload. Keep the prompts, tools and grading process the same and inspect invalid outputs as well as successful ones.
Artifact grading and repeatability study: What this evaluation measured
Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.
The final judgment comparison gives each reducer the same mapper findings. Defect finding and grade repeatability measure different aspects of review quality.
The source uses hand-labeled artifacts to test whether a grader catches actual defects instead of merely accepting an agent-produced audit script.
Artifact grading and repeatability study: Published results
| openai/gpt-5.6-luna | 83.1% | 19 of 23 |
| minimax/minimax-m3 | 82.5% | 16 of 23 |
No matching rows. Clear the filter to see all records.
23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table. Published 2026-08-08. Values retain their original units and grading scale.
Artifact grading and repeatability study: Methodology and sample
These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.
Harness
Map/reduce artifact grader versus independently hand-labeled defects
Corpus
23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table
Metric
Defect recall and repeated-grade agreement
Artifact grading and repeatability study: What the results do not establish
A reducer cannot discover a defect absent from its mapper inputs. These results do not establish end-to-end recall for every arbitrary artifact.
GPT 5.6 Luna and Tencent Hy3 tie on reported defect recall (83.1%). Their repeated-grade counts differ by one in 23 tests; that is not evidence of a universal review advantage.
This is a task-specific snapshot. Current grader defaults and newer model versions are not inferred from an August publication.