Build the smallest test you can explain
Run this example without an API key
Save the example
Open evaluation-example.mjs and save it into an empty folder. It contains one authored case and three fixed responses; it makes no network request.
Run it with Node.js
Run the command below. The included assertion checks the complete result.
Inspect the two failures
One response fits the schema but guesses that it can finish. The other is not JSON. Both stay in the denominator.
node evaluation-example.mjsThe output you should see
These are fixed teaching responses, not measured model performance. The narrow check verifies that the next action asks for a period. It does not judge the wording of the question or the quality of a financial brief.
{"fixture":"missing-fiscal-period","attempts":3,"schemaValid":2,"correct":1,"failures":2}Write the case before looking at answers
This is the authored prompt in the download, not a customer conversation or a prompt copied from the historical bakeoff. Its expected decision is ask. Keep the case, expected behavior and output contract together. Add a paired case with a specified period so a model that always asks cannot pass your whole suite.
Capture the real contract
For your own application, freeze the system instruction, message prefix, tool allowlist, schema, prompt version and referenced context versions. A prompt version alone misses changes to interpolated tool context.
Define failures explicitly
Record malformed JSON, schema violations, unavailable tools, invented resource IDs and provider failures separately. Schema validity does not establish a correct decision.
Prepare an NVDA financial brief. I have not chosen a fiscal period. Return JSON with decision (ask or finish) and question. Ask for the missing period before making financial claims.Move from a toy case to your application
NexusTrade's next-action replay rebuilds an unfinished agent decision from frozen messages and a frozen decision request. Candidate tools are not executed. The source checks schema validity, allowed tools, IDs present in the prefix, dependent tool batches and identical repeated actions before judge scoring.
Pin the model that actually served each request, the fixture hash, prompt and context versions, sample number, elapsed milliseconds, token usage and billing source. A substituted provider model is a separate arm, not another observation of the requested model. Keep failed attempts in a separate reliability report even when you cannot assign them a judge score.
Choose a fixed sample count and budget before running paid calls. Repeat the same case for every arm, persist each completed attempt, and resume only when the source hashes and settings still match. Run your captured source model as a fidelity control before comparing replacements.
Calibrate the judge before trusting its scores
Label examples independently, including known wrong actions that still look plausible. Give the judge the exact rubric and necessary evidence. Compare its decisions against those labels, repeat grading to measure stability, and regrade a shared sample with a judge from another model family.
A summary that clips evidence can turn a real number into a false hallucination finding. NexusTrade's grading code marks clipped prefix evidence rather than treating missing text as proof the model invented it. If the required answer uses a retired tool, count the case as unscorable; do not rewrite the expected action into a successor tool and call that accuracy.
What the published bakeoff actually established
The August 8, 2026 next-action study replayed 69 decisions three times across a 23-model campaign; 12 model rows were published. It measured individual decisions, not fresh complete trading runs. The tested input range was 73,397 to 83,363 tokens.
The historical Luna row reported score 89.2, schema validity 99.0%, $1.67 per 1,000 decisions and production p50 5.6s. The previous DeepSeek row reported 75.3, 98.1%, $6.23 and 17.9s. These are historical billed costs and production medians, not current API quotes or replay wall times.
Changing the judge narrowed their reported score gap from 13.9 to 1.5 points. Keep that uncertainty visible when choosing a deployment. A higher raw score alone does not settle the cost, latency and reliability tradeoff.
A release checklist you can use
Freeze and hold out
Version the cases and reserve unseen tasks before tuning prompts. Report what you changed and which population each score covers.
Report separate measures
Keep deterministic checks, calibrated judge scores, provider billing, cache accounting and latency separate. If cached-token usage was not retained, say so.
Verify the whole workflow separately
Replay quality is a screening result. Verify tool execution, permissions, artifact quality and recovery in a separately approved end-to-end test before deployment.