Shared agent run

AI Model Bake-Off: 10 Models Compete for the Best Trading Strategy

A comprehensive experiment comparing 10 AI models (GPT, Claude, Gemini, etc.) to find the best trading strategy. Includes backtest results, out-of-sample testing, and a leaderboard of the most profitable and risk-adjusted AI-generated portfolios.

Open interactive viewStart a new chat

Read the conversation below. Open the interactive view for charts, attachments, and continuing this run.

User

Use GPT-5.2, Opus 4.6, Sonnet 4.6, Gemini 3 Pro, Gemini 3.1 Pro, GPT5-mini, Gemini 3 Flash, Minimax 2.5, GLM-5, and Kimi K2.5 to independently try to create the best trading strategy in terms of raw gains and risk-adjusted returns. In each subagent, create the strategy based on data before 01/01/2024. Which model produces the consistently best strategies after 15 iterations?

Read the results and see which portfolios are winners. Then, we will perform a test-set. Based on the subagent data that you read, predict who the winner will be from 01/01/2024 to today. Were you right or wrong?\

Aurora

This is a fantastic experiment — a Model Bake-Off to determine which AI model produces the best trading strategies! We'll have 10 subagents, each powered by a different model, independently explore and create the best trading strategy they can find using data before 01/01/2024. Then we'll evaluate their results and run an out-of-sample test from 01/01/2024 to today.


The Core Question

Which AI model, given the same task and iteration budget, produces the most consistently profitable and risk-adjusted trading strategies?


Experiment Design

Phase 1: Training Period (Pre-2024 Data)
Each subagent will independently:

  • Define competing hypotheses about what drives returns
  • Create 4-8 portfolio variations testing different approaches
  • Backtest across multiple periods before 01/01/2024 (e.g., 2020, 2021, 2022, 2023)
  • Analyze results, iterate, and converge on their single best strategy
  • Each subagent gets 15 iterations to explore, test, and refine

Phase 2: Evaluation & Prediction
After all subagents complete, I will:

  • Read all subagent results and rank the winning portfolios by Total Return and Sortino Ratio
  • Present the top 2-3 strategies from each model
  • Make a prediction about which model's strategy will perform best on the unseen test set (01/01/2024 to today)

Phase 3: Out-of-Sample Test (01/01/2024 – Today)

  • Backtest the winning strategy from each model on the test set (01/01/2024 to 02/23/2026)
  • Compare actual results to my prediction
  • Declare the winner!

Step 1: Spawn 10 Subagents (Parallel Exploration)

Each subagent is a strategyExplorer with 15 max iterations, tasked with creating the best trading strategy using only pre-2024 data.

Subagent Model Model ID Cost
Subagent 1 GPT-5.2 azure/openai-responses/gpt-5.2-codex 9.25 tokens
Subagent 2 Claude Opus 4.6 anthropic/claude-opus-4.6 16.75 tokens
Subagent 3 Claude Sonnet 4.6 anthropic/claude-sonnet-4.6 10 tokens
Subagent 4 Gemini 3 Pro google/gemini-3.1-pro-preview 8 tokens
Subagent 5 Gemini 3.1 Pro google/gemini-3.1-pro-preview 8 tokens
Subagent 6 GPT-5-mini openai-responses/gpt-5-mini 1.25 tokens
Subagent 7 Gemini 3 Flash google/gemini-3-flash-preview 2 tokens
Subagent 8 MiniMax M2.5 minimaxi/MiniMax-M2.5 0.75 tokens
Subagent 9 GLM-5 zai/GLM-5 2.25 tokens
Subagent 10 Kimi K2.5 nebius/moonshotai/kimi-k2.5 1.75 tokens

Each subagent's mission:

"Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (use periods like 2020-2023). After exploring multiple approaches, converge on your single best strategy. Optimize for both total return and Sortino ratio."

Note: Gemini 3 Pro and Gemini 3.1 Pro are actually the same model ID (google/gemini-3.1-pro-preview), so Subagent 4 and 5 will both use it — this gives us a natural reproducibility check to see if the same model converges on similar strategies.


Step 2: Read All Subagent Results

  • Tool: Read Backtest (for each subagent's winning portfolios)
  • Compile a leaderboard ranking all 10 models by their best strategy's:
    • Total Return (training period)
    • Sortino Ratio (training period)
    • Maximum Drawdown (training period)
    • Consistency across multiple backtest periods

Step 3: Make Prediction & Run Out-of-Sample Test

  • Based on the training-period results, I will predict the winner on the test set (01/01/2024 – 02/23/2026)
  • Recreate each model's winning strategy and backtest it on the unseen period
  • Tools: Create YAML Portfolios, Backtest Portfolios, Read Backtest

Step 4: Final Analysis — Was I Right?

  • Compare my prediction to actual out-of-sample results
  • Analyze which strategies generalized well vs. which overfit to training data
  • Declare the overall winner of the Model Bake-Off
  • Tool: Read Backtest, then present final leaderboard

Summary

Step Action Tool(s)
1 Spawn 10 subagents (one per model) to independently explore and create the best strategy using pre-2024 data createSubagents (10 strategyExplorers)
2 Read all subagent results, compile leaderboard, identify each model's best strategy Read Backtest
3 Predict the test-set winner, then backtest all winning strategies on 01/01/2024 – 02/23/2026 Create YAML Portfolios, Backtest Portfolios
4 Read test-set results, compare to prediction, declare the winner Read Backtest

This is going to be an incredible experiment. Let's see which AI model is the best stock strategist! 🏆

User

Launching 10 subagents to explore strategies in parallel

User

Waiting for subagents to complete...

User

Waiting for subagents to complete...

Aurora

All 10 subagents complete.

Subagent 1: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: completed
  • Summary:
    US Stock Strategy Analysis: 11 Iterations Complete

Designed and tested 5 distinct trading logics (Momentum, Mean Reversion, Regime Adaptive, GARP, and Benchmark) over the 2020-2023 market cycle to identify the best risk-adjusted strategy.

Approach

Tested competing hypotheses: whether to buy strength (Momentum), buy weakness (Mean Reversion), or time the market (Regime Adaptive). Created 5 initial portfolios and ran 3 rounds of genetic optimization on the winner.

Key Discoveries

The Winner: 'Quality Momentum' was the only strategy to beat the SPY benchmark. It achieved +68.44% Return (vs SPY +53.97%) with a 0.79 Sortino Ratio. The logic is simple: Buy S&P 500 companies with positive Net Income and the highest 30-day Price Rate of Change.

What Failed:

  • Mean Reversion & GARP: Both significantly underperformed (+34% and +32%). In the post-2020 volatility, "cheap" stocks often got cheaper.
  • Regime Adaptive: While it reduced drawdown slightly, it missed the massive market rebounds, resulting in poor total returns (+28%).
  • Optimization: Genetic optimization attempts produced strategies with massive returns (+215% to +788%) but catastrophic risk (63-69% drawdowns). The algorithms "reverse overfitted" to aggressive volatility plays.

Top Performer

Quality Momentum is the robust winner. It balances capturing trends with a quality filter (Net Income > 0) to avoid speculative junk.

Next Steps

The Quality Momentum strategy is ready for deployment or further refinement. Future work could focus on testing different sell signals (e.g., trailing stops) to reduce the 35% drawdown without sacrificing the core momentum edge, as genetic optimization proved too aggressive.

Subagent 2: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: completed
  • Summary:
    US Stock Trading Strategy Optimization: Complete

Objective

Create the best possible trading strategy for US stocks optimized for total return and Sortino ratio, backtested on 2020-2023 data.

Methodology

Tested 5 competing hypotheses across 7 portfolios: Pure Momentum, Quality Momentum, Multi-Factor (quality + value), Regime-Adaptive Quality Momentum, and Mean Reversion Quality. All strategies were backtested from January 2020 to December 2023 (covering COVID crash, 2021 bull, 2022 bear, and 2023 recovery). The top two performers were then optimized using genetic algorithms.

Key Discoveries

Quality Filters Are Essential: Pure Momentum (39.57% return) underperformed SPY (53.97%), but adding a positive earnings filter transformed it into Quality Momentum (103.57% return)—a 161% improvement. Momentum without quality buys "momentum traps" that crash.

Contrarian Momentum Beats Trend-Following: The optimized strategy buys stocks with LOW 30-day ROC (recent laggards), not leaders. This contrarian approach captured mean reversion while the quality filter provided fundamental safety.

Mean Reversion Is Regime-Dependent: Mean Reversion Quality dominated the full-period backtest (132.32% return, 1.334 Sortino) but failed train/validation optimization splits. It excels in V-shaped recoveries (2020 COVID) but struggles in sustained downtrends (2022 bear market).

Value Filters Hurt Returns: The P/E < 25 filter in Multi-Factor was too restrictive, excluding growth winners. Return dropped to 60.54% vs 103.57% for Quality Momentum.

Top Performers

  1. Quality Momentum (Optimized) - 146.50% validation return, 1.10 Sharpe, 2.24 Sortino, 53.79% max drawdown
  2. Mean Reversion Quality - 132.32% return, 1.334 Sortino (full period only)
  3. Regime-Adaptive Quality Momentum - 82.08% return, 1.082 Sortino, 23.62% max drawdown (best risk control)

Final Recommendation

Deploy the Optimized Quality Momentum strategy with risk overlays (regime filter or position sizing limits). The strategy delivers exceptional returns with strong risk-adjusted metrics, though the 53.79% max drawdown requires active risk management. For risk-averse investors, the Regime-Adaptive Quality Momentum offers 82% return with only 23.62% drawdown.

Next Steps

  • Add 200-day SMA regime filter to reduce drawdowns
  • Test on 2024 forward data for out-of-sample validation
  • Consider blending with Regime-Adaptive for better risk-adjusted returns

Subagent 3: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: completed
  • Summary:
    US Stock Strategy Optimization: 6 Iterations Complete

Designed and tested four competing hypotheses to identify the most effective multi-factor trading strategy for US stocks during the volatile 2020–2023 period. The goal was to maximize total return and Sortino ratio while managing drawdowns across diverse market regimes.

Methodology

We tested four distinct approaches against a SPY Benchmark: Quality Momentum, Value-Quality (GARP), Regime-Adaptive Momentum (using a 200-day SMA filter), and Oversold Quality (Mean Reversion). All strategies were backtested from January 1, 2020, to December 31, 2023, using monthly rebalancing and a 10-stock concentration.

Key Discoveries

What Worked: Mean reversion significantly outperformed momentum. The 'Oversold Quality' strategy, which buys S&P 500 members with positive net income at 14-day RSI lows, was the only strategy to beat the benchmark, achieving a 55.01% return. This suggests that in the post-2020 market, "buying the dip" on blue-chip stocks was more effective than chasing trends.

What Failed: Momentum-based strategies ('Quality Momentum' and 'Regime-Adaptive') were catastrophic, with drawdowns exceeding 68% and 92% respectively. The rapid reversals in 2020 and 2022 made 90-day momentum signals too laggy, leading to significant capital erosion. Even the regime-switching logic failed to provide protection, as the 200-day SMA was frequently whipsawed.

Optimization Insights: A genetic optimization was performed on the 'Oversold Quality' portfolio. However, the results showed severe overfitting; while training returns were high, validation performance collapsed to -13.21%. The original parameters (14-day RSI) proved to be the most robust and generalizable.

Quantitative Highlights

  1. Oversold Quality: 55.01% Return, 0.69 Sortino, -44.81% Max Drawdown.
  2. SPY Benchmark: 53.97% Return, 0.74 Sortino, -33.33% Max Drawdown.
  3. Value-Quality: 19.67% Return, 1.28 Sortino, -83.32% Max Drawdown.

Next Steps

While 'Oversold Quality' is the clear winner, its 44% drawdown is higher than the benchmark. Future work should focus on adding a volatility-based position sizing rule or a secondary macro filter to reduce drawdown without sacrificing the alpha generated by the mean-reversion logic. Alternatively, we could re-run the optimization with a significantly larger population and generation count to find a more robust parameter set.

Subagent 4: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: completed
  • Summary:
    US Stock Trading Strategy Analysis Complete: Passive QQQ Wins

Objective: Find the optimal US stock trading strategy maximizing total return and Sortino ratio using 2020-2023 data only.

Research Approach: Tested 7 portfolios across 4 market regimes (COVID 2020, Bull 2021, Bear 2022, Recovery 2023): 2 benchmarks (SPY, QQQ) and 5 active strategies (Quality Momentum, Multi-Factor, Mean Reversion + Quality, Regime-Adaptive, Relative Strength Breakout). Optimized the best-performing active strategy (Mean Reversion) using genetic optimization.

Key Findings:

  • Winner: QQQ Benchmark — +92.64% return, 1.37 Sortino, nearly double SPY's +51.49%
  • Mean Reversion showed promise only in 2020 COVID recovery (+124.45%, 3.75 Sortino) but failed 2021-2023
  • Optimization Overfitting: Genetic optimization produced +3,011% training return but -59.51% validation return (117% Sharpe collapse)
  • Active Strategies Failed: Multi-Factor lost 86%, Relative Strength had zero trades, Quality Momentum and Regime-Adaptive underperformed

Conclusion: Simple buy-and-hold QQQ is the optimal strategy. Complex active strategies cannot consistently outperform a passive tech-heavy index fund. Mean Reversion shows potential but requires regime filtering to avoid overfitting.

Subagent 5: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: completed
  • Summary:
    US Stock Strategy Optimization: Complete Analysis (2020-2023)

Objective & Approach

User requested the best possible active trading strategy for US stocks using pre-2024 data, optimizing for total return and Sortino ratio. Tested four competing hypotheses against benchmarks across four distinct market regimes: COVID crash/recovery (2020), post-recovery bull (2021), bear market (2022), and new bull emergence (2023).

Portfolios Tested

Created and backtested 6 core strategies (with duplicate versions): SPY Benchmark, QQQ Benchmark, Quality Momentum (S&P 500 + TTM earnings filter + 90-day momentum), Regime-Adaptive Multi-Factor (200-day SMA regime switching), Deep Value Reversion (RSI < 30 + low P/E), and Composite Factor Ranking (40% momentum/30% value/30% quality weighting).

Key Discoveries

Passive vs Active: QQQ Benchmark dominated with 80.2% total return and 1.37 Sortino, though with severe -32.6% drawdown in 2022. Among active strategies, Quality Momentum emerged as the clear winner with 56.3% return and 0.72 Sortino, beating SPY by 4.2% while providing downside protection versus QQQ (-22.0% vs -32.6% in 2022).

Regime Performance: Quality Momentum excelled in crisis/recovery (+70.7% in 2020) but underperformed in sustained bull markets (+3.6% in 2021, +13.1% in 2023). The quality filter (positive TTM earnings) successfully prevented momentum crashes seen in unfiltered approaches.

Failed Approaches: Regime-Adaptive Multi-Factor produced -15.4% returns with negative Sortino (-0.17)—the elegant theory of switching between aggressive and defensive modes based on 200-day SMA failed due to whipsaws and missed the 2023 recovery entirely. Deep Value Reversion fell into value traps (-7.4% return), catching falling knives during the 2020 crash and 2023 banking crisis. Composite Factor Ranking suffered from factor dilution (45.8% return, inconsistent 0.56 Sortino).

Optimization Failure & Critical Insight

Attempted genetic algorithm optimization on Quality Momentum (training 2020-2022, validation 2023) but the process overfit severely to bull-market conditions. The "optimized" parameters (17 stocks, 13 max positions, >30 day triggers) delivered 1.60 Sharpe in training but collapsed to 0.40 in validation—a 75% degradation. This validated that the original domain-expert parameters (90-day lookback, monthly rebalance, 10 positions) generalize better than machine-optimized variants.

Final Recommendation

Deploy Quality Momentum with original parameters: S&P 500 universe, positive TTM earnings filter, top 10 by 90-day price momentum, equal-weighted, monthly rebalancing. This strategy provides the optimal active risk-adjusted return balance, exceptional crisis alpha, and superior downside protection versus tech-heavy benchmarks, while avoiding the overfitting pitfalls of optimized parameter sets.

Subagent 6: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: completed
  • Summary:

Ultimate Trading Strategy Discovery: Complete Analysis (2020-2023)

User's Original Request

Create the best possible trading strategy for US stocks with full creative freedom, optimized for total return and Sortino ratio, using only pre-2024 backtest data (2020-2023). Converge on a single best strategy through systematic testing and optimization.


Research Methodology

Approach: Tested 8 competing strategy hypotheses across 4 distinct market regimes (COVID crash 2020, bull market 2021, bear market 2022, recovery 2023) to identify the optimal balance between returns and risk-adjusted performance.

Hypotheses Tested:

  1. Adaptive Momentum + Quality (2 variants): Quality filters + momentum with regime-based allocation
  2. Multi-Factor (2 variants): Pure momentum vs. quality value
  3. Mean Reversion + Trend Filter: Oversold stocks in uptrends
  4. Concentrated Momentum: Top 5 stocks by composite momentum
  5. Multi-Strategy Ensemble: 40% momentum + 30% mean reversion + 30% quality value

Total Testing: 8 portfolios × 4 time periods = 32 backtests, plus 1 genetic optimization run (36 parameter combinations tested)


Key Discoveries

What Worked:

  • Quality Value emerged as the clear winner with 100% consistency (profitable in all 4 periods)
  • Only strategy to gain money in 2022 bear market (+1.06% vs SPY -18.19%)
  • Quality filters (positive earnings + ROE > 10%) combined with value discipline (low P/E) provided exceptional downside protection
  • 36.5% cumulative return with only 20.23% max drawdown

What Failed:

  • Pure momentum strategies collapsed in bear markets (Pure Momentum: -20.36% in 2022, -11.73% in 2023)
  • Mean reversion generated whipsaw losses (Trend-Filtered Reversion: -17.3% cumulative, worst performer)
  • Regime-switching complexity added no value (Regime-Adaptive optimization produced 68% drawdown)
  • Ensemble approaches underperformed single-factor strategies (17% cumulative vs 36.5% for Quality Value)

Unexpected Insights:

  • High-Conviction Top 5 achieved highest absolute return (95.8%) but failed in 2023 with 33% drawdown—concentration risk too high
  • Genetic optimization failed spectacularly: All optimized variants showed severe overfitting (training performance didn't translate to validation)
  • Simplicity won: The straightforward Quality Value strategy outperformed complex multi-strategy and regime-switching approaches

Performance Comparison (4-Year Results)

Strategy Cumulative Return 2022 Bear Market Win Rate Max Drawdown Verdict
Quality Value 36.5% +1.06% 100% 20.23% ✅ WINNER
High-Conviction Top 5 95.8% -11.84% 75% 33.03% ⚠️ Too volatile
Regime-Adaptive QM 72.0% -12.56% 75% 19.80% ❌ Optimization failed
SPY Benchmark 54.8% -18.19% 75% 33.33% Baseline
Quality Momentum Base 76.5% -18.63% 75% 31.68% Inconsistent
Ensemble Strategy 17.0% -19.45% 50% 38.31% Underperformed
Pure Momentum 16.8% -20.36% 50% 36.68% Failed
Trend-Filtered Reversion -17.3% -14.10% 25% 39.83% Worst

Final Recommendation: Quality Value Strategy

Exact Specifications:

  • Universe: S&P 500 stocks only
  • Quality Filters: Positive TTM net income AND ROE > 10%
  • Selection: Top 20 stocks by lowest P/E ratio (TTM)
  • Weighting: Equal weight (5% per position)
  • Rebalancing: Monthly

Why This Strategy Wins:

  1. ✅ Only strategy profitable in ALL 4 test periods (100% win rate)
  2. ✅ Made money in 2022 bear market (+1.06% vs SPY -18.19%)
  3. ✅ Best risk-adjusted profile (0.82 avg Sortino, 1.75 in 2023)
  4. ✅ Controlled drawdowns (20.23% vs 33.33% for SPY)
  5. ✅ Logically sound (fundamental quality + value discipline)

Trade-off: Sacrifices absolute return potential (36.5% vs 95.8% for High-Conviction) in exchange for perfect consistency, bear market protection, and lower stress.


Critical Lessons Learned

Optimization Failure: The genetic optimization of Regime-Adaptive Quality Momentum produced severe overfitting. Training performance (159% return) collapsed in validation (13% return with 68% drawdown). Small population size (6 individuals × 6 generations = 36 tests) and training period mismatch (excluded 2022 bear market) caused the algorithm to learn noise instead of signal.

Momentum Vulnerability: All pure momentum strategies failed in 2022-2023, losing 20-30% cumulatively. Without quality filters, momentum captures both rising stars and falling knives.

Complexity Penalty: Multi-strategy ensemble and regime-switching approaches underperformed simpler single-factor strategies, suggesting diversification across strategy types doesn't compensate for weak individual components.


Current State

COMPLETE: Systematic testing finished. Quality Value identified as the single best strategy meeting all success criteria (consistency, bear market resilience, acceptable drawdowns, logical parameters). Strategy is ready for deployment with exact specifications documented.

Success Criteria Met:

  • ✅ Consistency: 4/4 profitable periods (exceeded 3/4 target)
  • ✅ Max Drawdown: 20.23% (well below 30% threshold)
  • ✅ 2022 Performance: +1.06% (far exceeded "outperform SPY" target)
  • ✅ Robustness: Fundamental quality + value logic is sound
  • ⚠️ Total Return: 36.5% (below 60% target, but acceptable given risk profile)
  • ⚠️ Sortino Ratio: 0.82 avg (below 1.5 target, but 1.75 in 2023)

Next Steps: Quality Value strategy is validated and ready for implementation. For investors seeking higher absolute returns with acceptable volatility, High-Conviction Top 5 (95.8% return, 33% drawdown) is a viable alternative.

Subagent 7: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.

  • Status: max_iterations_reached
  • Summary:

Objective & Approach

User asked for the single best US-stock s

Message shortened in this preview.

User

Backtest the following portfolios for the period from January 1, 2024, to February 23, 2026, with $10,000 initial cash: 'Quality Momentum' (699c35da05fefd8f9953f5a2), 'Optimized Quality Momentum' (699c49f405fefd8f99592001), 'Oversold Quality' (699c363205fefd8f99540a59), 'QQQ Benchmark' (699c361b05fefd8f995404c5), 'Quality Momentum' (699c392805fefd8f9954adee), 'Quality Value' (699c360305fefd8f9953ff2d), 'ValueQuality_SP500_Top15' (699c35e705fefd8f9953f8be), 'Multi-Factor Concentrated - Top10' (699c365b05fefd8f995415fe), 'Quality Value' (699c3e5705fefd8f99558c04), and 'Refined QM Top 15 20d' (699c386c05fefd8f99548dce).

Aurora

I'll create the backtest configurations for the 10 winning strategies identified by our model subagents. This "test-set" backtest will run from January 1, 2024, to today (February 23, 2026) to determine which model's logic truly generalizes best to unseen data.

Prediction:
Based on the training data, Subagent 9 (GLM-5) produced the highest raw returns (501%), but the logic appears potentially overfit. Subagent 6 (GPT-5-mini) produced the most robust strategy, 'Quality Value', which was the only one profitable in every single training year (including the 2022 bear market). However, given the strong tech-led rally in 2024 and 2025, I predict Subagent 4 (Gemini 3 Pro) with the 'QQQ Benchmark' or Subagent 10 (Kimi K2.5) with 'Refined QM Top 15 20d' will likely take the lead in raw gains, while Subagent 6 will likely win on risk-adjusted returns (Sortino).

User

Analyze the backtest results for the following 10 portfolios for the period from January 1, 2024, to February 23, 2026: 'Quality Momentum' (699c35da05fefd8f9953f5a2), 'Optimized Quality Momentum' (699c49f405fefd8f99592001), 'Oversold Quality' (699c363205fefd8f99540a59), 'QQQ Benchmark' (699c361b05fefd8f995404c5), 'Quality Momentum' (699c392805fefd8f9954adee), 'Quality Value' (699c360305fefd8f9953ff2d), 'ValueQuality_SP500_Top15' (699c35e705fefd8f9953f8be), 'Multi-Factor Concentrated - Top10' (699c365b05fefd8f995415fe), 'Quality Value' (699c3e5705fefd8f99558c04), and 'Refined QM Top 15 20d' (699c386c05fefd8f99548dce). Please provide a leaderboard ranking them by Percent Change and Sortino Ratio, and compare their performance to the SPY benchmark.

Aurora

The results of the "Model Bake-Off" are in. This out-of-sample test (January 1, 2024, to February 23, 2026) reveals how well the strategies developed by the different AI models generalized to unseen market conditions.

1. Backtest Summary and Comparison

The baseline for this period is SPY, which returned 45.05% with a Sortino Ratio of 1.63 and a Max Drawdown of 20.19%.

Portfolio Name % Change Sortino Ratio Max Drawdown vs. SPY (Return)
Multi-Factor Concentrated - Top10 52.73% 1.46 32.51% Outperform
Quality Momentum (Subagent 1) 50.91% 1.31 21.98% Outperform
QQQ Benchmark 49.48% 1.42 23.84% Outperform
Refined QM Top 15 20d 49.42% 1.42 19.74% Outperform
Quality Value (Subagent 6) 47.98% 1.68 20.72% Outperform
Quality Momentum (Subagent 5) 40.69% 0.80 43.88% Underperform
Quality Value (Subagent 9) 37.79% 1.34 19.85% Underperform
ValueQuality_SP500_Top15 37.16% 1.27 23.07% Underperform
Oversold Quality 8.39% 0.41 29.08% Underperform
Optimized Quality Momentum -73.84% 15.38* 98.26% Collapse

*Note: The Sortino Ratio for the Optimized Quality Momentum is a statistical anomaly caused by the extreme magnitude of the loss and subsequent volatility calculation; the strategy effectively failed.


2. Trends

  • The Overfitting Trap: The "Optimized Quality Momentum" strategy, which showed the most aggressive gains during the training period, suffered a catastrophic collapse (-73.84%). This highlights the danger of genetic optimization without strict walk-forward validation.
  • Quality is the Anchor: Strategies that utilized "Quality" filters (Positive Net Income/ROE) generally stayed competitive. The Quality Value strategy from Subagent 6 was the only active strategy to beat SPY on a risk-adjusted basis (Sortino 1.68 vs 1.63).
  • Momentum vs. Drawdown: While Multi-Factor Concentrated and Quality Momentum delivered the highest raw gains, they did so with significantly higher drawdowns (up to 32.5%) compared to the SPY baseline.
  • Benchmark Resilience: SPY and QQQ remain incredibly difficult to beat. Only 5 out of 10 custom strategies managed to outperform SPY in raw returns, and only one beat it on a risk-adjusted (Sortino) basis.

3. Summary of Results

The winner in terms of raw gains is Subagent 8 (MiniMax 2.5) with the Multi-Factor Concentrated - Top10 strategy (+52.73%). However, the winner for risk-adjusted returns is Subagent 6 (GPT-5-mini) with the Quality Value strategy, which achieved a Sortino of 1.68.

The most balanced performer was Refined QM Top 15 20d, which managed to beat SPY's returns while actually maintaining a lower maximum drawdown (19.74% vs 20.19%).


4. Leaderboard & Prediction Review

Leaderboard by Percent Change:

  1. Multi-Factor Concentrated - Top10 (52.73%) 🏆
  2. Quality Momentum (Subagent 1) (50.91%)
  3. QQQ Benchmark (49.48%)
  4. Refined QM Top 15 20d (49.42%)
  5. Quality Value (Subagent 6) (47.98%)

Leaderboard by Sortino Ratio:

  1. Quality Value (Subagent 6) (1.68) 🏆
  2. Multi-Factor Concentrated - Top10 (1.46)
  3. QQQ Benchmark (1.42)
  4. Refined QM Top 15 20d (1.42)
  5. Quality Value (Subagent 9) (1.34)

Were you right or wrong?
I was partially right. I predicted that QQQ or Refined QM would lead in raw gains; they came in 3rd and 4th, narrowly beaten by the Multi-Factor and Quality Momentum strategies. I also predicted that Subagent 6 (Quality Value) would win on risk-adjusted returns, which was correct (Sortino 1.68).

Portfolio Suggestion for Optimization:
I suggest optimizing Refined QM Top 15 20d. It has already proven it can beat SPY's returns while maintaining a lower drawdown. If a genetic optimizer can refine its entry/exit triggers to improve the Sortino ratio from 1.42 to >1.70, it would be the definitive "Alpha" strategy of this group.

User

🏆 The Model Bake-Off: Final Results & Analysis

This experiment pitted 10 different AI models against each other to see which could design the most robust trading strategy using only historical data (pre-2024). We then tested these 'winners' against the live market from January 1, 2024, to February 23, 2026.

1. The Winners' Circle

Rank Model Strategy % Change Sortino Result
1st (Gains) MiniMax 2.5 Multi-Factor Concentrated 52.73% 1.46 Raw Returns Winner
1st (Risk) GPT-5-mini Quality Value 47.98% 1.68 Risk-Adjusted Winner
Baseline Market SPY Benchmark 45.05% 1.63 Market Baseline

2. Detailed Leaderboard (Out-of-Sample Test Set)

  1. Multi-Factor Concentrated (Subagent 8 - MiniMax 2.5): 52.73% Return | 1.46 Sortino. This strategy used a complex blend of momentum, volatility, and fundamental quality to edge out the market.
  2. Quality Momentum (Subagent 1 - GPT-5.2): 50.91% Return | 1.31 Sortino. A strong runner-up that proved simple trend-following with a quality filter is highly effective.
  3. QQQ Benchmark (Subagent 4 - Gemini 3 Pro): 49.48% Return | 1.42 Sortino. Proved that beating the tech index is extremely difficult for active strategies.
  4. Refined QM Top 15 20d (Subagent 10 - Kimi K2.5): 49.42% Return | 1.42 Sortino. Notable for having the lowest Maximum Drawdown (19.74%) of all winners.
  5. Quality Value (Subagent 6 - GPT-5-mini): 47.98% Return | 1.68 Sortino. The only active strategy to beat SPY on a risk-adjusted basis.

3. Prediction Review: Was I Right?

  • Prediction for Gains: I predicted the QQQ Benchmark or Refined QM would lead. I was partially right; they performed excellently (3rd and 4th place), but were narrowly beaten by the aggressive Multi-Factor and Quality Momentum strategies.
  • Prediction for Risk-Adjusted: I predicted Subagent 6 (Quality Value) would win on Sortino Ratio. This was correct (1.68 Sortino vs SPY's 1.63).

4. Key Takeaway: The Overfitting Trap

The most important lesson from this bake-off came from Subagent 2 (Claude Opus 4.6). Its 'Optimized Quality Momentum' strategy had the most impressive backtests in the training period but collapsed by -73.84% in the test set. This confirms that genetic optimization without strict out-of-sample validation often results in strategies that 'memorize' the past rather than 'predicting' the future.

Final Verdict

While MiniMax 2.5 produced the highest raw gains, GPT-5-mini produced the most 'investable' strategy (Quality Value), as it was the only one to deliver higher returns than the S&P 500 while also providing better risk-adjusted performance. For a balanced approach, Kimi K2.5's Refined QM strategy is the most robust, offering market-beating returns with the lowest volatility.

Aurora

🏆 The Model Bake-Off: Final Results & Analysis

This experiment pitted 10 different AI models against each other to see which could design the most robust trading strategy using only historical data (pre-2024). We then tested these 'winners' against the live market from January 1, 2024, to February 23, 2026.

1. The Winners' Circle

Rank Model Strategy % Change Sortino Result
1st (Gains) MiniMax 2.5 Multi-Factor Concentrated 52.73% 1.46 Raw Returns Winner
1st (Risk) GPT-5-mini Quality Value 47.98% 1.68 Risk-Adjusted Winner
Baseline Market SPY Benchmark 45.05% 1.63 Market Baseline

2. Detailed Leaderboard (Out-of-Sample Test Set)

  1. Multi-Factor Concentrated (Subagent 8 - MiniMax 2.5): 52.73% Return | 1.46 Sortino. This strategy used a complex blend of momentum, volatility, and fundamental quality to edge out the market.
  2. Quality Momentum (Subagent 1 - GPT-5.2): 50.91% Return | 1.31 Sortino. A strong runner-up that proved simple trend-following with a quality filter is highly effective.
  3. QQQ Benchmark (Subagent 4 - Gemini 3 Pro): 49.48% Return | 1.42 Sortino. Proved that beating the tech index is extremely difficult for active strategies.
  4. Refined QM Top 15 20d (Subagent 10 - Kimi K2.5): 49.42% Return | 1.42 Sortino. Notable for having the lowest Maximum Drawdown (19.74%) of all winners.
  5. Quality Value (Subagent 6 - GPT-5-mini): 47.98% Return | 1.68 Sortino. The only active strategy to beat SPY on a risk-adjusted basis.

3. Prediction Review: Was I Right?

  • Prediction for Gains: I predicted the QQQ Benchmark or Refined QM would lead. I was partially right; they performed excellently (3rd and 4th place), but were narrowly beaten by the aggressive Multi-Factor and Quality Momentum strategies.
  • Prediction for Risk-Adjusted: I predicted Subagent 6 (Quality Value) would win on Sortino Ratio. This was correct (1.68 Sortino vs SPY's 1.63).

4. Key Takeaway: The Overfitting Trap

The most important lesson from this bake-off came from Subagent 2 (Claude Opus 4.6). Its 'Optimized Quality Momentum' strategy had the most impressive backtests in the training period but collapsed by -73.84% in the test set. This confirms that genetic optimization without strict out-of-sample validation often results in strategies that 'memorize' the past rather than 'predicting' the future.

Final Verdict

While MiniMax 2.5 produced the highest raw gains, GPT-5-mini produced the most 'investable' strategy (Quality Value), as it was the only one to deliver higher returns than the S&P 500 while also providing better risk-adjusted performance. For a balanced approach, Kimi K2.5's Refined QM strategy is the most robust, offering market-beating returns with the lowest volatility.

This is a bounded preview of a long transcript. Open the interactive view for the complete conversation.

Open interactive viewStart a new chat