Published model evaluations

Qwen3.7 Flash: Trading Task Evaluations

See how qwen/qwen3.7-flash performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
qwen/qwen3.7-flash
Published studies
1
Provider
qwen

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

The provider name identifies the model family. The tools and harness around the model determine what it can do.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Five models completing one sandbox research task

Qwen3.7 Flash recorded D → D → D for grade trajectory in the 2026-08-08 publication. 5 models on the same single research task; real Modal virtual machines.

Study-wide context: Luna reached B+ three times and had the lowest reported completed-run wall time and cost. The oscillating grades and single-task sample remain material limitations.

This is one task, not a broad coding benchmark. A favorable result cannot be generalized to all trading research code.

The early remediation loop oscillated between grades. Later diagnoses found flaws in agent-written audit code; a successful execution status alone did not establish the correctness of delivered analysis.

Flash Lite failed out after 24 minutes and has no reported total cost in the table. Missing cost is not represented as zero.

Grade trajectoryD → D → D
Wall124m
Cost$2.85

Continue exploring