Published model evaluations

Hy3: Trading Task Evaluations

See how tencent/hy3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
tencent/hy3
Published studies
1
Provider
tencent

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

The provider name identifies the model family. The tools and harness around the model determine what it can do.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Artifact grading and repeatability study

Hy3 recorded 83.1% for defects found in the 2026-08-08 publication. 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table.

Study-wide context: Luna and Hy3 share the reported recall. Luna repeated its grade on 19/23 reruns, Hy3 on 18/23 and MiniMax M3 on 16/23.

A reducer cannot discover a defect absent from its mapper inputs. These results do not establish end-to-end recall for every arbitrary artifact.

GPT 5.6 Luna and Tencent Hy3 tie on reported defect recall (83.1%). Their repeated-grade counts differ by one in 23 tests; that is not evidence of a universal review advantage.

This is a task-specific snapshot. Current grader defaults and newer model versions are not inferred from an August publication.

Defects found83.1%
Same grade on re-run18 of 23

Continue exploring