Published model evaluations

Best AI Models for Agent Execution: Published Evaluations

Which models performed well on agent execution in NexusTrade's published evaluations? Compare the task, exact versions, costs and evidence limitations.

As of 2026-08-08

What the agent execution evidence supports

As published on 2026-08-08: The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.

Test against your own workload

Check whether the tested task and required output match your workload. Use a fixed set of representative examples, repeat the runs and inspect failures. Compare cost and latency alongside quality.

These are published historical results. Internal customer conversations remain private. The results do not establish the current production default or promise profitable trading.

What this evaluation measured

The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.

Models choose the next action inside the same ReAct agent loop. The reported context window in this replay was 73,397 to 83,363 input tokens per decision.

The table reports mean grader score alongside deterministic schema checks, provider billing and production median latency. These columns should be evaluated separately.

Published results

openai/gpt-5.6-luna-pro91.7100.0%$25.059.9s
openai/gpt-5.6-luna89.299.0%$1.675.6s
meta/muse-spark-1.286.199.5%$44.237.2s
x-ai/grok-build-0.184.7100.0%$26.8815.4s
google/gemini-3-flash-preview83.397.1%$24.956.7s
z-ai/glm-5.282.498.1%$25.3715.7s
google/gemini-3.6-flash80.1100.0%$24.217.7s
deepseek/deepseek-v4-flash75.398.1%$6.2317.9s
mistralai/mistral-small-260373.097.6%$1.006.3s
google/gemini-3.5-flash-lite72.399.0%$7.172.0s
nvidia/nemotron-3-ultra-550b57.177.5%$35.28n/a
poolside/laguna-xs-2.151.166.5%$5.584.3s

23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows. Published 2026-08-08. Values retain their original units and grading scale.

Methodology and sample

These reviewed results come from an evaluation that has already been published. Private conversation traces are excluded. Opening this page does not run a new benchmark.

  1. Harness

    NexusTrade Agent V6 next-action replay

  2. Corpus

    23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows

  3. Metric

    Mean next-action score; schema validity, billed cost and production p50 are separate

What the results do not establish

Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.

Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.

Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.

Read the original evidence

Continue exploring