Published model evaluations

Muse Spark 1.2: Trading Task Evaluations

See how meta/muse-spark-1.2 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
meta/muse-spark-1.2
Published studies
1
Provider
meta

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

Muse Spark is the underlying model evaluated here. Meta Muse is a separate personal-agent product with its own tools and runtime.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Next-action execution model bakeoff

Muse Spark 1.2 recorded 86.1 for mean score in the 2026-08-08 publication. 23-model campaign; 69 frozen decisions × 3 samples; 12 published model rows.

Study-wide context: The published deployment decision selected GPT 5.6 Luna because it improved score, schema validity, billed cost and latency versus the old DeepSeek default. GPT 5.6 Luna Pro had a higher raw score and a substantially higher cost.

Only 12 of the 23 campaign arms are in the public table. The catalog does not fabricate the unpublished rows or their failure counts.

Cross-vendor regrading changed the reported gaps. In the source, Luna versus DeepSeek moved from a 13.9-point score gap to 1.5 points under another judge. Quality claims remain judge-dependent.

Reported costs benefited from reused prompt prefixes. The source did not retain the cached-token split. Historical billed cost is not a current API price quote.

Mean score86.1
Schema valid99.5%
$ / 1k decisions$44.23
Prod p507.2s

Continue exploring