Model evaluations
Artifact grading and repeatability study
Published 2026-08-08: 23 hand-labeled artifacts covering 14 task types; 3 reducer models in the public table. Defect recall and repeated-grade agreement. See the results, how they were measured and where the findings stop.
Explore →Best AI models for Agent Execution: published evaluations
Which models performed well on agent execution in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.
Explore →Best AI models for Agent Planning: published evaluations
Which models performed well on agent planning in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.
Explore →Best AI models for algorithmic trading: evidence by task
Compare published model evaluations for options strategies, planning, agent execution, artifact review, sandbox code and stock SQL screening. Each task keeps its own methodology.
Explore →Best AI models for Artifact Review: published evaluations
Which models performed well on artifact review in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.
Explore →Best AI models for Options Strategy Development: published evaluations
Which models performed well on options strategy development in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.
Explore →Best AI models for Sandbox Code Generation: published evaluations
Which models performed well on sandbox code generation in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.
Explore →Best AI models for Stock Screening: published evaluations
Which models performed well on stock screening in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.
Explore →Claude Opus 5: trading task evaluations
See how anthropic/claude-opus-5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Claude Sonnet 4.6: trading task evaluations
See how anthropic/claude-sonnet-4.6 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Claude Sonnet 5: trading task evaluations
See how anthropic/claude-sonnet-5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Deepseek V4 Flash: trading task evaluations
See how deepseek/deepseek-v4-flash performed in 3 trading task evaluations, with exact version, harness, dates and limitations.
Explore →Deepseek V4 Pro: trading task evaluations
See how deepseek/deepseek-v4-pro performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Eleven models building an options strategy
Published 2026-04-06: 11 model runs; same task, account context, tools and 25-iteration budget. Goal-oriented evaluator score / 100. See the results, how they were measured and where the findings stop.
Explore →Five models completing one sandbox research task
Published 2026-08-08: 5 models on the same single research task; real Modal virtual machines. Grade trajectory, elapsed wall time and billed run cost. See the results, how they were measured and where the findings stop.
Explore →Gemini 3 Flash Preview vs Claude Sonnet 4.6: options-strategy development
Compare exact model versions on 1 shared published evaluation of options-strategy development. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →Gemini 3 Flash Preview vs GPT-5.4: options-strategy development
Compare exact model versions on 1 shared published evaluation of options-strategy development. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →Gemini 3 Flash Preview vs Grok 4.20: options-strategy development
Compare exact model versions on 1 shared published evaluation of options-strategy development. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →Gemini 3 Flash Preview: trading task evaluations
See how google/gemini-3-flash-preview performed in 3 trading task evaluations, with exact version, harness, dates and limitations.
Explore →Gemini 3 Pro Preview: trading task evaluations
See how google/gemini-3-pro-preview performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Gemini 3.1 Flash Lite: trading task evaluations
See how google/gemini-3.1-flash-lite performed in 2 trading task evaluations, with exact version, harness, dates and limitations.
Explore →Gemini 3.5 Flash Lite: trading task evaluations
See how google/gemini-3.5-flash-lite performed in 3 trading task evaluations, with exact version, harness, dates and limitations.
Explore →Gemini 3.5 Flash: trading task evaluations
See how google/gemini-3.5-flash performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Gemini 3.6 Flash: trading task evaluations
See how google/gemini-3.6-flash performed in 3 trading task evaluations, with exact version, harness, dates and limitations.
Explore →Gemma 4 31B: trading task evaluations
See how google/gemma-4-31b performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →GLM 5.2: trading task evaluations
See how z-ai/glm-5.2 performed in 3 trading task evaluations, with exact version, harness, dates and limitations.
Explore →GPT 5 Mini: trading task evaluations
See how openai/gpt-5-mini performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →GPT 5.6 Luna Pro: trading task evaluations
See how openai/gpt-5.6-luna-pro performed in 2 trading task evaluations, with exact version, harness, dates and limitations.
Explore →GPT 5.6 Luna vs Claude Opus 5: agent planning
Compare exact model versions on 1 shared published evaluation of agent planning. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs Claude Sonnet 5: agent planning
Compare exact model versions on 1 shared published evaluation of agent planning. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs Deepseek V4 Flash: planning, execution and sandbox research
Compare exact model versions on 3 shared published evaluations of planning, execution and sandbox research. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs Gemini 3 Flash Preview: agent execution and stock-screening SQL
Compare exact model versions on 2 shared published evaluations of agent execution and stock-screening SQL. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs Gemini 3.6 Flash: planning, execution and stock-screening SQL
Compare exact model versions on 3 shared published evaluations of planning, execution and stock-screening SQL. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs GLM 5.2: planning, execution and sandbox research
Compare exact model versions on 3 shared published evaluations of planning, execution and sandbox research. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs MiniMax M3: artifact review
Compare exact model versions on 1 shared published evaluation of artifact review. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs Muse Spark 1.1: agent planning
Compare exact model versions on 1 shared published evaluation of agent planning. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna vs Muse Spark 1.2: agent execution
Compare exact model versions on 1 shared published evaluation of agent execution. See the observed tradeoffs, rubric, sample, date and cost units.
Explore →GPT 5.6 Luna: trading task evaluations
See how openai/gpt-5.6-luna performed in 5 trading task evaluations, with exact version, harness, dates and limitations.
Explore →GPT 5.6 Terra Pro: trading task evaluations
See how openai/gpt-5.6-terra-pro performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →GPT-5.4-mini: trading task evaluations
See how openai/gpt-5.4-mini performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →GPT-5.4: trading task evaluations
See how openai/gpt-5.4 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Grok 4.20: trading task evaluations
See how x-ai/grok-4.20 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Grok 4.5: trading task evaluations
See how x-ai/grok-4.5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Grok Build 0.1: trading task evaluations
See how x-ai/grok-build-0.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Hy3: trading task evaluations
See how tencent/hy3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Kimi K2.5: trading task evaluations
See how moonshotai/kimi-k2.5 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Kimi K3: trading task evaluations
See how moonshotai/kimi-k3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Laguna Xs 2.1: trading task evaluations
See how poolside/laguna-xs-2.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →MiMo v2 Pro: trading task evaluations
See how xiaomi/mimo-v2-pro performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →MiniMax M3: trading task evaluations
See how minimax/minimax-m3 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Mistral Small 2603: trading task evaluations
See how mistralai/mistral-small-2603 performed in 2 trading task evaluations, with exact version, harness, dates and limitations.
Explore →Models answering 22 stock questions with SQL
Published 2026-08-08: 22 natural-language questions; 6 published model rows. Average answer score, success rate and average execution time. See the results, how they were measured and where the findings stop.
Explore →Muse Spark 1.1: trading task evaluations
See how meta/muse-spark-1.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Muse Spark 1.2: trading task evaluations
See how meta/muse-spark-1.2 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Nemotron 3 Ultra 550B: trading task evaluations
See how nvidia/nemotron-3-ultra-550b performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Next-action execution model bakeoff
Published 2026-08-08: 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Mean next-action score; schema validity, billed cost and production p50 are separate. See the results, how they were measured and where the findings stop.
Explore →Planning models on 29 frozen conversations
Published 2026-08-08: 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table. Mean graded plan score; standard error reported separately. See the results, how they were measured and where the findings stop.
Explore →Qwen3.7 Flash: trading task evaluations
See how qwen/qwen3.7-flash performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →Qwen3.8 Max: trading task evaluations
See how qwen/qwen3.8-max performed in 1 trading task evaluation, with exact version, harness, dates and limitations.
Explore →