Published model evaluations

Muse Spark 1.1: Trading Task Evaluations

See how meta/muse-spark-1.1 performed in 1 trading task evaluation, with exact version, harness, dates and limitations.

As of 2026-08-08

Exact tested ID
meta/muse-spark-1.1
Published studies
1
Provider
meta

Evidence for this exact version

The results apply to the exact model ID above. Newer versions need their own evaluation. Scores with different grading scales are kept separate, and a model result alone does not evaluate a complete trading agent.

Muse Spark is the underlying model evaluated here. Meta Muse is a separate personal-agent product with its own tools and runtime.

Use these results to choose a test

Start with the study closest to your task. Compare this version with the other models in that study using the same rubric and cost units. A strong result on planning does not establish stock-screening or coding quality.

Before adopting a version, try representative inputs in your own harness. Check failure cases and valid outputs alongside score, elapsed time and billed cost. The tables below describe the published runs; they are not current price quotes or live-return forecasts.

Planning models on 29 frozen conversations

Muse Spark 1.1 recorded 0.759 for score in the 2026-08-08 publication. 16-model campaign, 29 frozen conversations; 14 model rows shown in the public table.

Study-wide context: GPT 5.6 Luna was the published cost-oriented selection among similar-quality planning candidates. Muse Spark 1.1 had the highest raw mean; those are different claims.

The public table omits two poolside arms for space. This snapshot contains the 14 published rows, not a reconstruction of their private outputs.

The top raw score is Muse Spark 1.1 at 0.759. GPT 5.6 Luna scored 0.731 at a reported $0.0012 per decision; the published selection favored cost among similar-quality candidates, rather than claiming the highest score.

The source reports a paired comparison with Opus as inconclusive. This page does not turn small raw gaps into a confident quality winner.

Score0.759
+/- SE0.079
$ / decision$0.0229
p509.6s

Continue exploring