Research library

Model evaluations

Compare published AI model results by task, exact version, grading method, cost and latency. Start with the benchmark that matches your workload.

Model evaluations

Artifact grading and repeatability study

Published 2026-08-08: 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table. Defect recall and repeated-grade agreement. See the results, how they were measured and where the findings stop.

Explore →

Best AI models for Agent Execution: published evaluations

Which models performed well on agent execution in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

Explore →

Best AI models for Agent Planning: published evaluations

Which models performed well on agent planning in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

Explore →

Best AI models for algorithmic trading: evidence by task

Compare published model evaluations for options strategies, planning, agent execution, artifact review, sandbox code and stock SQL screening. Each task keeps its own methodology.

Explore →

Best AI models for Artifact Review: published evaluations

Which models performed well on artifact review in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

Explore →

Best AI models for Options Strategy Development: published evaluations

Which models performed well on options strategy development in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

Explore →

Best AI models for Sandbox Code Generation: published evaluations

Which models performed well on sandbox code generation in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

Explore →

Best AI models for Stock Screening: published evaluations

Which models performed well on stock screening in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

Explore →

Claude Opus 5: trading task evaluations

See how anthropic/claude-opus-5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Claude Sonnet 4.6: trading task evaluations

See how anthropic/claude-sonnet-4.6 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Claude Sonnet 5: trading task evaluations

See how anthropic/claude-sonnet-5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Deepseek V4 Flash: trading task evaluations

See how deepseek/deepseek-v4-flash performed in 3 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

Deepseek V4 Pro: trading task evaluations

See how deepseek/deepseek-v4-pro performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Eleven models building an options strategy

Published 2026-04-06: 11 model runs; same task, account context, tools and 25-iteration budget. Goal-oriented evaluator score / 100. See the results, how they were measured and where the findings stop.

Explore →

Five models completing one sandbox research task

Published 2026-08-08: 5 models on the same single research task; real Modal virtual machines. Grade trajectory, elapsed wall time and billed run cost. See the results, how they were measured and where the findings stop.

Explore →

Gemini 3 Flash Preview vs Claude Sonnet 4.6: options-strategy development

Compare exact model versions on 1 shared published evaluation of options-strategy development. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

Gemini 3 Flash Preview vs GPT-5.4: options-strategy development

Compare exact model versions on 1 shared published evaluation of options-strategy development. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

Gemini 3 Flash Preview vs Grok 4.20: options-strategy development

Compare exact model versions on 1 shared published evaluation of options-strategy development. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

Gemini 3 Flash Preview: trading task evaluations

See how google/gemini-3-flash-preview performed in 3 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

Gemini 3 Pro Preview: trading task evaluations

See how google/gemini-3-pro-preview performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Gemini 3.1 Flash Lite: trading task evaluations

See how google/gemini-3.1-flash-lite performed in 2 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

Gemini 3.5 Flash Lite: trading task evaluations

See how google/gemini-3.5-flash-lite performed in 3 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

Gemini 3.5 Flash: trading task evaluations

See how google/gemini-3.5-flash performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Gemini 3.6 Flash: trading task evaluations

See how google/gemini-3.6-flash performed in 3 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

Gemma 4 31B: trading task evaluations

See how google/gemma-4-31b performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

GLM 5.2: trading task evaluations

See how z-ai/glm-5.2 performed in 3 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

GPT 5 Mini: trading task evaluations

See how openai/gpt-5-mini performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

GPT 5.6 Luna Pro: trading task evaluations

See how openai/gpt-5.6-luna-pro performed in 2 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

GPT 5.6 Luna vs Claude Opus 5: agent planning

Compare exact model versions on 1 shared published evaluation of agent planning. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs Claude Sonnet 5: agent planning

Compare exact model versions on 1 shared published evaluation of agent planning. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs Deepseek V4 Flash: planning, execution and sandbox research

Compare exact model versions on 3 shared published evaluations of planning, execution and sandbox research. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs Gemini 3 Flash Preview: agent execution and stock-screening SQL

Compare exact model versions on 2 shared published evaluations of agent execution and stock-screening SQL. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs Gemini 3.6 Flash: planning, execution and stock-screening SQL

Compare exact model versions on 3 shared published evaluations of planning, execution and stock-screening SQL. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs GLM 5.2: planning, execution and sandbox research

Compare exact model versions on 3 shared published evaluations of planning, execution and sandbox research. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs MiniMax M3: artifact review

Compare exact model versions on 1 shared published evaluation of artifact review. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs Muse Spark 1.1: agent planning

Compare exact model versions on 1 shared published evaluation of agent planning. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna vs Muse Spark 1.2: agent execution

Compare exact model versions on 1 shared published evaluation of agent execution. See the observed tradeoffs, rubric, sample, date and cost units.

Explore →

GPT 5.6 Luna: trading task evaluations

See how openai/gpt-5.6-luna performed in 5 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

GPT 5.6 Terra Pro: trading task evaluations

See how openai/gpt-5.6-terra-pro performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

GPT-5.4-mini: trading task evaluations

See how openai/gpt-5.4-mini performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

GPT-5.4: trading task evaluations

See how openai/gpt-5.4 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Grok 4.20: trading task evaluations

See how x-ai/grok-4.20 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Grok 4.5: trading task evaluations

See how x-ai/grok-4.5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Grok Build 0.1: trading task evaluations

See how x-ai/grok-build-0.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Hy3: trading task evaluations

See how tencent/hy3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Kimi K2.5: trading task evaluations

See how moonshotai/kimi-k2.5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Kimi K3: trading task evaluations

See how moonshotai/kimi-k3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Laguna Xs 2.1: trading task evaluations

See how poolside/laguna-xs-2.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

MiMo v2 Pro: trading task evaluations

See how xiaomi/mimo-v2-pro performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

MiniMax M3: trading task evaluations

See how minimax/minimax-m3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Mistral Small 2603: trading task evaluations

See how mistralai/mistral-small-2603 performed in 2 trading task evaluations, with exact version, harness, dates and limitations.

Explore →

Models answering 22 stock questions with SQL

Published 2026-08-08: 22 natural-language questions; 6 published model rows. Average answer score, success rate and average execution time. See the results, how they were measured and where the findings stop.

Explore →

Muse Spark 1.1: trading task evaluations

See how meta/muse-spark-1.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Muse Spark 1.2: trading task evaluations

See how meta/muse-spark-1.2 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Nemotron 3 Ultra 550B: trading task evaluations

See how nvidia/nemotron-3-ultra-550b performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Next-action execution model bakeoff

Published 2026-08-08: 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Mean next-action score; schema validity, billed cost and production p50 are separate. See the results, how they were measured and where the findings stop.

Explore →

Planning models on 29 frozen conversations

Published 2026-08-08: 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Mean graded plan score; standard error reported separately. See the results, how they were measured and where the findings stop.

Explore →

Qwen3.7 Flash: trading task evaluations

See how qwen/qwen3.7-flash performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Qwen3.8 Max: trading task evaluations

See how qwen/qwen3.8-max performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

Explore →

Continue exploring