Published model evaluations

Best AI Models for Artifact Review: Published Evaluations

Which models performed well on artifact review in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

As of 2026-08-08

What the artifact review evidence supports

As published on 2026-08-08: Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.

Test against your own workload

Check whether the tested task and required output match your workload. Use a fixed set of representative examples, repeat the runs and inspect failures. Compare cost and latency alongside quality.

These are published historical results. Internal customer conversations remain private. The results do not establish the current production default or promise profitable trading.

What this evaluation measured

Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.

The final judgment comparison gives each reducer the same mapper findings. Defect finding and grade repeatability measure different aspects of review quality.

The source uses hand-labeled artifacts to test whether a grader catches actual defects instead of merely accepting an agent-produced audit script.

Published results

openai/gpt-5.6-luna83.1%19 of 23
tencent/hy383.1%18 of 23
minimax/minimax-m382.5%16 of 23

23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table. Published 2026-08-08. Values retain their original units and grading scale.

Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    Map/reduce artifact grader versus independently hand-labeled defects

  2. Corpus

    23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table

  3. Metric

    Defect recall and repeated-grade agreement

What the results do not establish

A reducer cannot discover a defect absent from its mapper inputs. These results do not establish end-to-end recall for every arbitrary artifact.

GPT 5.6 Luna and Tencent Hy3 tie on reported defect recall (83.1%). Their repeated-grade counts differ by one in 23 tests; that is not evidence of a universal review advantage.

This is a task-specific snapshot. Current grader defaults and newer model versions are not inferred from an August publication.

Read the original evidence

Continue exploring