Evidence for this exact version
The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.
The provider name identifies the model family. The tools and harness around the model determine what it can do.
Use these results to choose a test
Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.
Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.
Artifact grading and repeatability study
Hy3 recorded 83.1% for defects found in the 2026-08-08 publication. 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table.
Study-wide context: Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.
A reducer cannot discover a defect absent from its mapper inputs. These results do not establish end-to-end recall for every arbitrary artifact.
GPT 5.6 Luna and Tencent Hy3 tie on reported defect recall (83.1%). Their repeated-grade counts differ by one in 23 tests; that is not evidence of a universal review advantage.
This is a task-specific snapshot. Current grader defaults and newer model versions are not inferred from an August publication.
| Defects found | 83.1% |
| Same grade on re-run | 18 of 23 |
No matching rows. Clear the filter to see all records.