Shared agent run
AI Model Bake-Off: 10 Models Compete for the Best Trading Strategy
A comprehensive experiment comparing 10 AI models (GPT, Claude, Gemini, etc.) to find the best trading strategy. Includes backtest results, out-of-sample testing, and a leaderboard of the most profitable and risk-adjusted AI-generated portfolios.
Read the conversation below. Open the interactive view for charts, attachments, and continuing this run.
User
Use GPT-5.2, Opus 4.6, Sonnet 4.6, Gemini 3 Pro, Gemini 3.1 Pro, GPT5-mini, Gemini 3 Flash, Minimax 2.5, GLM-5, and Kimi K2.5 to independently try to create the best trading strategy in terms of raw gains and risk-adjusted returns. In each subagent, create the strategy based on data before 01/01/2024. Which model produces the consistently best strategies after 15 iterations?
Read the results and see which portfolios are winners. Then, we will perform a test-set. Based on the subagent data that you read, predict who the winner will be from 01/01/2024 to today. Were you right or wrong?\
Aurora
This is a fantastic experiment — a Model Bake-Off to determine which AI model produces the best trading strategies! We'll have 10 subagents, each powered by a different model, independently explore and create the best trading strategy they can find using data before 01/01/2024. Then we'll evaluate their results and run an out-of-sample test from 01/01/2024 to today.
The Core Question
Which AI model, given the same task and iteration budget, produces the most consistently profitable and risk-adjusted trading strategies?
Experiment Design
Phase 1: Training Period (Pre-2024 Data)
Each subagent will independently:
- Define competing hypotheses about what drives returns
- Create 4-8 portfolio variations testing different approaches
- Backtest across multiple periods before 01/01/2024 (e.g., 2020, 2021, 2022, 2023)
- Analyze results, iterate, and converge on their single best strategy
- Each subagent gets 15 iterations to explore, test, and refine
Phase 2: Evaluation & Prediction
After all subagents complete, I will:
- Read all subagent results and rank the winning portfolios by Total Return and Sortino Ratio
- Present the top 2-3 strategies from each model
- Make a prediction about which model's strategy will perform best on the unseen test set (01/01/2024 to today)
Phase 3: Out-of-Sample Test (01/01/2024 – Today)
- Backtest the winning strategy from each model on the test set (01/01/2024 to 02/23/2026)
- Compare actual results to my prediction
- Declare the winner!
Step 1: Spawn 10 Subagents (Parallel Exploration)
Each subagent is a strategyExplorer with 15 max iterations, tasked with creating the best trading strategy using only pre-2024 data.
| Subagent | Model | Model ID | Cost |
|---|---|---|---|
| Subagent 1 | GPT-5.2 | azure/openai-responses/gpt-5.2-codex |
9.25 tokens |
| Subagent 2 | Claude Opus 4.6 | anthropic/claude-opus-4.6 |
16.75 tokens |
| Subagent 3 | Claude Sonnet 4.6 | anthropic/claude-sonnet-4.6 |
10 tokens |
| Subagent 4 | Gemini 3 Pro | google/gemini-3.1-pro-preview |
8 tokens |
| Subagent 5 | Gemini 3.1 Pro | google/gemini-3.1-pro-preview |
8 tokens |
| Subagent 6 | GPT-5-mini | openai-responses/gpt-5-mini |
1.25 tokens |
| Subagent 7 | Gemini 3 Flash | google/gemini-3-flash-preview |
2 tokens |
| Subagent 8 | MiniMax M2.5 | minimaxi/MiniMax-M2.5 |
0.75 tokens |
| Subagent 9 | GLM-5 | zai/GLM-5 |
2.25 tokens |
| Subagent 10 | Kimi K2.5 | nebius/moonshotai/kimi-k2.5 |
1.75 tokens |
Each subagent's mission:
"Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (use periods like 2020-2023). After exploring multiple approaches, converge on your single best strategy. Optimize for both total return and Sortino ratio."
Note: Gemini 3 Pro and Gemini 3.1 Pro are actually the same model ID (google/gemini-3.1-pro-preview), so Subagent 4 and 5 will both use it — this gives us a natural reproducibility check to see if the same model converges on similar strategies.
Step 2: Read All Subagent Results
- Tool:
Read Backtest(for each subagent's winning portfolios) - Compile a leaderboard ranking all 10 models by their best strategy's:
- Total Return (training period)
- Sortino Ratio (training period)
- Maximum Drawdown (training period)
- Consistency across multiple backtest periods
Step 3: Make Prediction & Run Out-of-Sample Test
- Based on the training-period results, I will predict the winner on the test set (01/01/2024 – 02/23/2026)
- Recreate each model's winning strategy and backtest it on the unseen period
- Tools:
Create YAML Portfolios,Backtest Portfolios,Read Backtest
Step 4: Final Analysis — Was I Right?
- Compare my prediction to actual out-of-sample results
- Analyze which strategies generalized well vs. which overfit to training data
- Declare the overall winner of the Model Bake-Off
- Tool:
Read Backtest, then present final leaderboard
Summary
| Step | Action | Tool(s) |
|---|---|---|
| 1 | Spawn 10 subagents (one per model) to independently explore and create the best strategy using pre-2024 data | createSubagents (10 strategyExplorers) |
| 2 | Read all subagent results, compile leaderboard, identify each model's best strategy | Read Backtest |
| 3 | Predict the test-set winner, then backtest all winning strategies on 01/01/2024 – 02/23/2026 | Create YAML Portfolios, Backtest Portfolios |
| 4 | Read test-set results, compare to prediction, declare the winner | Read Backtest |
This is going to be an incredible experiment. Let's see which AI model is the best stock strategist! 🏆
User
Launching 10 subagents to explore strategies in parallel
User
Waiting for subagents to complete...
User
Waiting for subagents to complete...
Aurora
All 10 subagents complete.
Subagent 1: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: completed
- Summary:
US Stock Strategy Analysis: 11 Iterations Complete
Designed and tested 5 distinct trading logics (Momentum, Mean Reversion, Regime Adaptive, GARP, and Benchmark) over the 2020-2023 market cycle to identify the best risk-adjusted strategy.
Approach
Tested competing hypotheses: whether to buy strength (Momentum), buy weakness (Mean Reversion), or time the market (Regime Adaptive). Created 5 initial portfolios and ran 3 rounds of genetic optimization on the winner.
Key Discoveries
The Winner: 'Quality Momentum' was the only strategy to beat the SPY benchmark. It achieved +68.44% Return (vs SPY +53.97%) with a 0.79 Sortino Ratio. The logic is simple: Buy S&P 500 companies with positive Net Income and the highest 30-day Price Rate of Change.
What Failed:
- Mean Reversion & GARP: Both significantly underperformed (+34% and +32%). In the post-2020 volatility, "cheap" stocks often got cheaper.
- Regime Adaptive: While it reduced drawdown slightly, it missed the massive market rebounds, resulting in poor total returns (+28%).
- Optimization: Genetic optimization attempts produced strategies with massive returns (+215% to +788%) but catastrophic risk (63-69% drawdowns). The algorithms "reverse overfitted" to aggressive volatility plays.
Top Performer
Quality Momentum is the robust winner. It balances capturing trends with a quality filter (Net Income > 0) to avoid speculative junk.
Next Steps
The Quality Momentum strategy is ready for deployment or further refinement. Future work could focus on testing different sell signals (e.g., trailing stops) to reduce the 35% drawdown without sacrificing the core momentum edge, as genetic optimization proved too aggressive.
Subagent 2: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: completed
- Summary:
US Stock Trading Strategy Optimization: Complete
Objective
Create the best possible trading strategy for US stocks optimized for total return and Sortino ratio, backtested on 2020-2023 data.
Methodology
Tested 5 competing hypotheses across 7 portfolios: Pure Momentum, Quality Momentum, Multi-Factor (quality + value), Regime-Adaptive Quality Momentum, and Mean Reversion Quality. All strategies were backtested from January 2020 to December 2023 (covering COVID crash, 2021 bull, 2022 bear, and 2023 recovery). The top two performers were then optimized using genetic algorithms.
Key Discoveries
Quality Filters Are Essential: Pure Momentum (39.57% return) underperformed SPY (53.97%), but adding a positive earnings filter transformed it into Quality Momentum (103.57% return)—a 161% improvement. Momentum without quality buys "momentum traps" that crash.
Contrarian Momentum Beats Trend-Following: The optimized strategy buys stocks with LOW 30-day ROC (recent laggards), not leaders. This contrarian approach captured mean reversion while the quality filter provided fundamental safety.
Mean Reversion Is Regime-Dependent: Mean Reversion Quality dominated the full-period backtest (132.32% return, 1.334 Sortino) but failed train/validation optimization splits. It excels in V-shaped recoveries (2020 COVID) but struggles in sustained downtrends (2022 bear market).
Value Filters Hurt Returns: The P/E < 25 filter in Multi-Factor was too restrictive, excluding growth winners. Return dropped to 60.54% vs 103.57% for Quality Momentum.
Top Performers
- Quality Momentum (Optimized) - 146.50% validation return, 1.10 Sharpe, 2.24 Sortino, 53.79% max drawdown
- Mean Reversion Quality - 132.32% return, 1.334 Sortino (full period only)
- Regime-Adaptive Quality Momentum - 82.08% return, 1.082 Sortino, 23.62% max drawdown (best risk control)
Final Recommendation
Deploy the Optimized Quality Momentum strategy with risk overlays (regime filter or position sizing limits). The strategy delivers exceptional returns with strong risk-adjusted metrics, though the 53.79% max drawdown requires active risk management. For risk-averse investors, the Regime-Adaptive Quality Momentum offers 82% return with only 23.62% drawdown.
Next Steps
- Add 200-day SMA regime filter to reduce drawdowns
- Test on 2024 forward data for out-of-sample validation
- Consider blending with Regime-Adaptive for better risk-adjusted returns
Subagent 3: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: completed
- Summary:
US Stock Strategy Optimization: 6 Iterations Complete
Designed and tested four competing hypotheses to identify the most effective multi-factor trading strategy for US stocks during the volatile 2020–2023 period. The goal was to maximize total return and Sortino ratio while managing drawdowns across diverse market regimes.
Methodology
We tested four distinct approaches against a SPY Benchmark: Quality Momentum, Value-Quality (GARP), Regime-Adaptive Momentum (using a 200-day SMA filter), and Oversold Quality (Mean Reversion). All strategies were backtested from January 1, 2020, to December 31, 2023, using monthly rebalancing and a 10-stock concentration.
Key Discoveries
What Worked: Mean reversion significantly outperformed momentum. The 'Oversold Quality' strategy, which buys S&P 500 members with positive net income at 14-day RSI lows, was the only strategy to beat the benchmark, achieving a 55.01% return. This suggests that in the post-2020 market, "buying the dip" on blue-chip stocks was more effective than chasing trends.
What Failed: Momentum-based strategies ('Quality Momentum' and 'Regime-Adaptive') were catastrophic, with drawdowns exceeding 68% and 92% respectively. The rapid reversals in 2020 and 2022 made 90-day momentum signals too laggy, leading to significant capital erosion. Even the regime-switching logic failed to provide protection, as the 200-day SMA was frequently whipsawed.
Optimization Insights: A genetic optimization was performed on the 'Oversold Quality' portfolio. However, the results showed severe overfitting; while training returns were high, validation performance collapsed to -13.21%. The original parameters (14-day RSI) proved to be the most robust and generalizable.
Quantitative Highlights
- Oversold Quality: 55.01% Return, 0.69 Sortino, -44.81% Max Drawdown.
- SPY Benchmark: 53.97% Return, 0.74 Sortino, -33.33% Max Drawdown.
- Value-Quality: 19.67% Return, 1.28 Sortino, -83.32% Max Drawdown.
Next Steps
While 'Oversold Quality' is the clear winner, its 44% drawdown is higher than the benchmark. Future work should focus on adding a volatility-based position sizing rule or a secondary macro filter to reduce drawdown without sacrificing the alpha generated by the mean-reversion logic. Alternatively, we could re-run the optimization with a significantly larger population and generation count to find a more robust parameter set.
Subagent 4: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: completed
- Summary:
US Stock Trading Strategy Analysis Complete: Passive QQQ Wins
Objective: Find the optimal US stock trading strategy maximizing total return and Sortino ratio using 2020-2023 data only.
Research Approach: Tested 7 portfolios across 4 market regimes (COVID 2020, Bull 2021, Bear 2022, Recovery 2023): 2 benchmarks (SPY, QQQ) and 5 active strategies (Quality Momentum, Multi-Factor, Mean Reversion + Quality, Regime-Adaptive, Relative Strength Breakout). Optimized the best-performing active strategy (Mean Reversion) using genetic optimization.
Key Findings:
- Winner: QQQ Benchmark — +92.64% return, 1.37 Sortino, nearly double SPY's +51.49%
- Mean Reversion showed promise only in 2020 COVID recovery (+124.45%, 3.75 Sortino) but failed 2021-2023
- Optimization Overfitting: Genetic optimization produced +3,011% training return but -59.51% validation return (117% Sharpe collapse)
- Active Strategies Failed: Multi-Factor lost 86%, Relative Strength had zero trades, Quality Momentum and Regime-Adaptive underperformed
Conclusion: Simple buy-and-hold QQQ is the optimal strategy. Complex active strategies cannot consistently outperform a passive tech-heavy index fund. Mean Reversion shows potential but requires regime filtering to avoid overfitting.
Subagent 5: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: completed
- Summary:
US Stock Strategy Optimization: Complete Analysis (2020-2023)
Objective & Approach
User requested the best possible active trading strategy for US stocks using pre-2024 data, optimizing for total return and Sortino ratio. Tested four competing hypotheses against benchmarks across four distinct market regimes: COVID crash/recovery (2020), post-recovery bull (2021), bear market (2022), and new bull emergence (2023).
Portfolios Tested
Created and backtested 6 core strategies (with duplicate versions): SPY Benchmark, QQQ Benchmark, Quality Momentum (S&P 500 + TTM earnings filter + 90-day momentum), Regime-Adaptive Multi-Factor (200-day SMA regime switching), Deep Value Reversion (RSI < 30 + low P/E), and Composite Factor Ranking (40% momentum/30% value/30% quality weighting).
Key Discoveries
Passive vs Active: QQQ Benchmark dominated with 80.2% total return and 1.37 Sortino, though with severe -32.6% drawdown in 2022. Among active strategies, Quality Momentum emerged as the clear winner with 56.3% return and 0.72 Sortino, beating SPY by 4.2% while providing downside protection versus QQQ (-22.0% vs -32.6% in 2022).
Regime Performance: Quality Momentum excelled in crisis/recovery (+70.7% in 2020) but underperformed in sustained bull markets (+3.6% in 2021, +13.1% in 2023). The quality filter (positive TTM earnings) successfully prevented momentum crashes seen in unfiltered approaches.
Failed Approaches: Regime-Adaptive Multi-Factor produced -15.4% returns with negative Sortino (-0.17)—the elegant theory of switching between aggressive and defensive modes based on 200-day SMA failed due to whipsaws and missed the 2023 recovery entirely. Deep Value Reversion fell into value traps (-7.4% return), catching falling knives during the 2020 crash and 2023 banking crisis. Composite Factor Ranking suffered from factor dilution (45.8% return, inconsistent 0.56 Sortino).
Optimization Failure & Critical Insight
Attempted genetic algorithm optimization on Quality Momentum (training 2020-2022, validation 2023) but the process overfit severely to bull-market conditions. The "optimized" parameters (17 stocks, 13 max positions, >30 day triggers) delivered 1.60 Sharpe in training but collapsed to 0.40 in validation—a 75% degradation. This validated that the original domain-expert parameters (90-day lookback, monthly rebalance, 10 positions) generalize better than machine-optimized variants.
Final Recommendation
Deploy Quality Momentum with original parameters: S&P 500 universe, positive TTM earnings filter, top 10 by 90-day price momentum, equal-weighted, monthly rebalancing. This strategy provides the optimal active risk-adjusted return balance, exceptional crisis alpha, and superior downside protection versus tech-heavy benchmarks, while avoiding the overfitting pitfalls of optimized parameter sets.
Subagent 6: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: completed
- Summary:
Ultimate Trading Strategy Discovery: Complete Analysis (2020-2023)
User's Original Request
Create the best possible trading strategy for US stocks with full creative freedom, optimized for total return and Sortino ratio, using only pre-2024 backtest data (2020-2023). Converge on a single best strategy through systematic testing and optimization.
Research Methodology
Approach: Tested 8 competing strategy hypotheses across 4 distinct market regimes (COVID crash 2020, bull market 2021, bear market 2022, recovery 2023) to identify the optimal balance between returns and risk-adjusted performance.
Hypotheses Tested:
- Adaptive Momentum + Quality (2 variants): Quality filters + momentum with regime-based allocation
- Multi-Factor (2 variants): Pure momentum vs. quality value
- Mean Reversion + Trend Filter: Oversold stocks in uptrends
- Concentrated Momentum: Top 5 stocks by composite momentum
- Multi-Strategy Ensemble: 40% momentum + 30% mean reversion + 30% quality value
Total Testing: 8 portfolios × 4 time periods = 32 backtests, plus 1 genetic optimization run (36 parameter combinations tested)
Key Discoveries
What Worked:
- Quality Value emerged as the clear winner with 100% consistency (profitable in all 4 periods)
- Only strategy to gain money in 2022 bear market (+1.06% vs SPY -18.19%)
- Quality filters (positive earnings + ROE > 10%) combined with value discipline (low P/E) provided exceptional downside protection
- 36.5% cumulative return with only 20.23% max drawdown
What Failed:
- Pure momentum strategies collapsed in bear markets (Pure Momentum: -20.36% in 2022, -11.73% in 2023)
- Mean reversion generated whipsaw losses (Trend-Filtered Reversion: -17.3% cumulative, worst performer)
- Regime-switching complexity added no value (Regime-Adaptive optimization produced 68% drawdown)
- Ensemble approaches underperformed single-factor strategies (17% cumulative vs 36.5% for Quality Value)
Unexpected Insights:
- High-Conviction Top 5 achieved highest absolute return (95.8%) but failed in 2023 with 33% drawdown—concentration risk too high
- Genetic optimization failed spectacularly: All optimized variants showed severe overfitting (training performance didn't translate to validation)
- Simplicity won: The straightforward Quality Value strategy outperformed complex multi-strategy and regime-switching approaches
Performance Comparison (4-Year Results)
| Strategy | Cumulative Return | 2022 Bear Market | Win Rate | Max Drawdown | Verdict |
|---|---|---|---|---|---|
| Quality Value | 36.5% | +1.06% | 100% | 20.23% | ✅ WINNER |
| High-Conviction Top 5 | 95.8% | -11.84% | 75% | 33.03% | ⚠️ Too volatile |
| Regime-Adaptive QM | 72.0% | -12.56% | 75% | 19.80% | ❌ Optimization failed |
| SPY Benchmark | 54.8% | -18.19% | 75% | 33.33% | Baseline |
| Quality Momentum Base | 76.5% | -18.63% | 75% | 31.68% | Inconsistent |
| Ensemble Strategy | 17.0% | -19.45% | 50% | 38.31% | Underperformed |
| Pure Momentum | 16.8% | -20.36% | 50% | 36.68% | Failed |
| Trend-Filtered Reversion | -17.3% | -14.10% | 25% | 39.83% | Worst |
Final Recommendation: Quality Value Strategy
Exact Specifications:
- Universe: S&P 500 stocks only
- Quality Filters: Positive TTM net income AND ROE > 10%
- Selection: Top 20 stocks by lowest P/E ratio (TTM)
- Weighting: Equal weight (5% per position)
- Rebalancing: Monthly
Why This Strategy Wins:
- ✅ Only strategy profitable in ALL 4 test periods (100% win rate)
- ✅ Made money in 2022 bear market (+1.06% vs SPY -18.19%)
- ✅ Best risk-adjusted profile (0.82 avg Sortino, 1.75 in 2023)
- ✅ Controlled drawdowns (20.23% vs 33.33% for SPY)
- ✅ Logically sound (fundamental quality + value discipline)
Trade-off: Sacrifices absolute return potential (36.5% vs 95.8% for High-Conviction) in exchange for perfect consistency, bear market protection, and lower stress.
Critical Lessons Learned
Optimization Failure: The genetic optimization of Regime-Adaptive Quality Momentum produced severe overfitting. Training performance (159% return) collapsed in validation (13% return with 68% drawdown). Small population size (6 individuals × 6 generations = 36 tests) and training period mismatch (excluded 2022 bear market) caused the algorithm to learn noise instead of signal.
Momentum Vulnerability: All pure momentum strategies failed in 2022-2023, losing 20-30% cumulatively. Without quality filters, momentum captures both rising stars and falling knives.
Complexity Penalty: Multi-strategy ensemble and regime-switching approaches underperformed simpler single-factor strategies, suggesting diversification across strategy types doesn't compensate for weak individual components.
Current State
COMPLETE: Systematic testing finished. Quality Value identified as the single best strategy meeting all success criteria (consistency, bear market resilience, acceptable drawdowns, logical parameters). Strategy is ready for deployment with exact specifications documented.
Success Criteria Met:
- ✅ Consistency: 4/4 profitable periods (exceeded 3/4 target)
- ✅ Max Drawdown: 20.23% (well below 30% threshold)
- ✅ 2022 Performance: +1.06% (far exceeded "outperform SPY" target)
- ✅ Robustness: Fundamental quality + value logic is sound
- ⚠️ Total Return: 36.5% (below 60% target, but acceptable given risk profile)
- ⚠️ Sortino Ratio: 0.82 avg (below 1.5 target, but 1.75 in 2023)
Next Steps: Quality Value strategy is validated and ready for implementation. For investors seeking higher absolute returns with acceptable volatility, High-Conviction Top 5 (95.8% return, 33% drawdown) is a viable alternative.
Subagent 7: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio.
- Status: max_iterations_reached
- Summary:
Objective & Approach
User asked for the single best US-stock s
Message shortened in this preview.
User
Backtest the following portfolios for the period from January 1, 2024, to February 23, 2026, with $10,000 initial cash: 'Quality Momentum' (699c35da05fefd8f9953f5a2), 'Optimized Quality Momentum' (699c49f405fefd8f99592001), 'Oversold Quality' (699c363205fefd8f99540a59), 'QQQ Benchmark' (699c361b05fefd8f995404c5), 'Quality Momentum' (699c392805fefd8f9954adee), 'Quality Value' (699c360305fefd8f9953ff2d), 'ValueQuality_SP500_Top15' (699c35e705fefd8f9953f8be), 'Multi-Factor Concentrated - Top10' (699c365b05fefd8f995415fe), 'Quality Value' (699c3e5705fefd8f99558c04), and 'Refined QM Top 15 20d' (699c386c05fefd8f99548dce).
Aurora
I'll create the backtest configurations for the 10 winning strategies identified by our model subagents. This "test-set" backtest will run from January 1, 2024, to today (February 23, 2026) to determine which model's logic truly generalizes best to unseen data.
Prediction:
Based on the training data, Subagent 9 (GLM-5) produced the highest raw returns (501%), but the logic appears potentially overfit. Subagent 6 (GPT-5-mini) produced the most robust strategy, 'Quality Value', which was the only one profitable in every single training year (including the 2022 bear market). However, given the strong tech-led rally in 2024 and 2025, I predict Subagent 4 (Gemini 3 Pro) with the 'QQQ Benchmark' or Subagent 10 (Kimi K2.5) with 'Refined QM Top 15 20d' will likely take the lead in raw gains, while Subagent 6 will likely win on risk-adjusted returns (Sortino).
User
Analyze the backtest results for the following 10 portfolios for the period from January 1, 2024, to February 23, 2026: 'Quality Momentum' (699c35da05fefd8f9953f5a2), 'Optimized Quality Momentum' (699c49f405fefd8f99592001), 'Oversold Quality' (699c363205fefd8f99540a59), 'QQQ Benchmark' (699c361b05fefd8f995404c5), 'Quality Momentum' (699c392805fefd8f9954adee), 'Quality Value' (699c360305fefd8f9953ff2d), 'ValueQuality_SP500_Top15' (699c35e705fefd8f9953f8be), 'Multi-Factor Concentrated - Top10' (699c365b05fefd8f995415fe), 'Quality Value' (699c3e5705fefd8f99558c04), and 'Refined QM Top 15 20d' (699c386c05fefd8f99548dce). Please provide a leaderboard ranking them by Percent Change and Sortino Ratio, and compare their performance to the SPY benchmark.
Aurora
The results of the "Model Bake-Off" are in. This out-of-sample test (January 1, 2024, to February 23, 2026) reveals how well the strategies developed by the different AI models generalized to unseen market conditions.
1. Backtest Summary and Comparison
The baseline for this period is SPY, which returned 45.05% with a Sortino Ratio of 1.63 and a Max Drawdown of 20.19%.
| Portfolio Name | % Change | Sortino Ratio | Max Drawdown | vs. SPY (Return) |
|---|---|---|---|---|
| Multi-Factor Concentrated - Top10 | 52.73% | 1.46 | 32.51% | Outperform |
| Quality Momentum (Subagent 1) | 50.91% | 1.31 | 21.98% | Outperform |
| QQQ Benchmark | 49.48% | 1.42 | 23.84% | Outperform |
| Refined QM Top 15 20d | 49.42% | 1.42 | 19.74% | Outperform |
| Quality Value (Subagent 6) | 47.98% | 1.68 | 20.72% | Outperform |
| Quality Momentum (Subagent 5) | 40.69% | 0.80 | 43.88% | Underperform |
| Quality Value (Subagent 9) | 37.79% | 1.34 | 19.85% | Underperform |
| ValueQuality_SP500_Top15 | 37.16% | 1.27 | 23.07% | Underperform |
| Oversold Quality | 8.39% | 0.41 | 29.08% | Underperform |
| Optimized Quality Momentum | -73.84% | 15.38* | 98.26% | Collapse |
*Note: The Sortino Ratio for the Optimized Quality Momentum is a statistical anomaly caused by the extreme magnitude of the loss and subsequent volatility calculation; the strategy effectively failed.
2. Trends
- The Overfitting Trap: The "Optimized Quality Momentum" strategy, which showed the most aggressive gains during the training period, suffered a catastrophic collapse (-73.84%). This highlights the danger of genetic optimization without strict walk-forward validation.
- Quality is the Anchor: Strategies that utilized "Quality" filters (Positive Net Income/ROE) generally stayed competitive. The Quality Value strategy from Subagent 6 was the only active strategy to beat SPY on a risk-adjusted basis (Sortino 1.68 vs 1.63).
- Momentum vs. Drawdown: While Multi-Factor Concentrated and Quality Momentum delivered the highest raw gains, they did so with significantly higher drawdowns (up to 32.5%) compared to the SPY baseline.
- Benchmark Resilience: SPY and QQQ remain incredibly difficult to beat. Only 5 out of 10 custom strategies managed to outperform SPY in raw returns, and only one beat it on a risk-adjusted (Sortino) basis.
3. Summary of Results
The winner in terms of raw gains is Subagent 8 (MiniMax 2.5) with the Multi-Factor Concentrated - Top10 strategy (+52.73%). However, the winner for risk-adjusted returns is Subagent 6 (GPT-5-mini) with the Quality Value strategy, which achieved a Sortino of 1.68.
The most balanced performer was Refined QM Top 15 20d, which managed to beat SPY's returns while actually maintaining a lower maximum drawdown (19.74% vs 20.19%).
4. Leaderboard & Prediction Review
Leaderboard by Percent Change:
- Multi-Factor Concentrated - Top10 (52.73%) 🏆
- Quality Momentum (Subagent 1) (50.91%)
- QQQ Benchmark (49.48%)
- Refined QM Top 15 20d (49.42%)
- Quality Value (Subagent 6) (47.98%)
Leaderboard by Sortino Ratio:
- Quality Value (Subagent 6) (1.68) 🏆
- Multi-Factor Concentrated - Top10 (1.46)
- QQQ Benchmark (1.42)
- Refined QM Top 15 20d (1.42)
- Quality Value (Subagent 9) (1.34)
Were you right or wrong?
I was partially right. I predicted that QQQ or Refined QM would lead in raw gains; they came in 3rd and 4th, narrowly beaten by the Multi-Factor and Quality Momentum strategies. I also predicted that Subagent 6 (Quality Value) would win on risk-adjusted returns, which was correct (Sortino 1.68).
Portfolio Suggestion for Optimization:
I suggest optimizing Refined QM Top 15 20d. It has already proven it can beat SPY's returns while maintaining a lower drawdown. If a genetic optimizer can refine its entry/exit triggers to improve the Sortino ratio from 1.42 to >1.70, it would be the definitive "Alpha" strategy of this group.
User
🏆 The Model Bake-Off: Final Results & Analysis
This experiment pitted 10 different AI models against each other to see which could design the most robust trading strategy using only historical data (pre-2024). We then tested these 'winners' against the live market from January 1, 2024, to February 23, 2026.
1. The Winners' Circle
| Rank | Model | Strategy | % Change | Sortino | Result |
|---|---|---|---|---|---|
| 1st (Gains) | MiniMax 2.5 | Multi-Factor Concentrated | 52.73% | 1.46 | Raw Returns Winner |
| 1st (Risk) | GPT-5-mini | Quality Value | 47.98% | 1.68 | Risk-Adjusted Winner |
| Baseline | Market | SPY Benchmark | 45.05% | 1.63 | Market Baseline |
2. Detailed Leaderboard (Out-of-Sample Test Set)
- Multi-Factor Concentrated (Subagent 8 - MiniMax 2.5): 52.73% Return | 1.46 Sortino. This strategy used a complex blend of momentum, volatility, and fundamental quality to edge out the market.
- Quality Momentum (Subagent 1 - GPT-5.2): 50.91% Return | 1.31 Sortino. A strong runner-up that proved simple trend-following with a quality filter is highly effective.
- QQQ Benchmark (Subagent 4 - Gemini 3 Pro): 49.48% Return | 1.42 Sortino. Proved that beating the tech index is extremely difficult for active strategies.
- Refined QM Top 15 20d (Subagent 10 - Kimi K2.5): 49.42% Return | 1.42 Sortino. Notable for having the lowest Maximum Drawdown (19.74%) of all winners.
- Quality Value (Subagent 6 - GPT-5-mini): 47.98% Return | 1.68 Sortino. The only active strategy to beat SPY on a risk-adjusted basis.
3. Prediction Review: Was I Right?
- Prediction for Gains: I predicted the QQQ Benchmark or Refined QM would lead. I was partially right; they performed excellently (3rd and 4th place), but were narrowly beaten by the aggressive Multi-Factor and Quality Momentum strategies.
- Prediction for Risk-Adjusted: I predicted Subagent 6 (Quality Value) would win on Sortino Ratio. This was correct (1.68 Sortino vs SPY's 1.63).
4. Key Takeaway: The Overfitting Trap
The most important lesson from this bake-off came from Subagent 2 (Claude Opus 4.6). Its 'Optimized Quality Momentum' strategy had the most impressive backtests in the training period but collapsed by -73.84% in the test set. This confirms that genetic optimization without strict out-of-sample validation often results in strategies that 'memorize' the past rather than 'predicting' the future.
Final Verdict
While MiniMax 2.5 produced the highest raw gains, GPT-5-mini produced the most 'investable' strategy (Quality Value), as it was the only one to deliver higher returns than the S&P 500 while also providing better risk-adjusted performance. For a balanced approach, Kimi K2.5's Refined QM strategy is the most robust, offering market-beating returns with the lowest volatility.
Aurora
🏆 The Model Bake-Off: Final Results & Analysis
This experiment pitted 10 different AI models against each other to see which could design the most robust trading strategy using only historical data (pre-2024). We then tested these 'winners' against the live market from January 1, 2024, to February 23, 2026.
1. The Winners' Circle
| Rank | Model | Strategy | % Change | Sortino | Result |
|---|---|---|---|---|---|
| 1st (Gains) | MiniMax 2.5 | Multi-Factor Concentrated | 52.73% | 1.46 | Raw Returns Winner |
| 1st (Risk) | GPT-5-mini | Quality Value | 47.98% | 1.68 | Risk-Adjusted Winner |
| Baseline | Market | SPY Benchmark | 45.05% | 1.63 | Market Baseline |
2. Detailed Leaderboard (Out-of-Sample Test Set)
- Multi-Factor Concentrated (Subagent 8 - MiniMax 2.5): 52.73% Return | 1.46 Sortino. This strategy used a complex blend of momentum, volatility, and fundamental quality to edge out the market.
- Quality Momentum (Subagent 1 - GPT-5.2): 50.91% Return | 1.31 Sortino. A strong runner-up that proved simple trend-following with a quality filter is highly effective.
- QQQ Benchmark (Subagent 4 - Gemini 3 Pro): 49.48% Return | 1.42 Sortino. Proved that beating the tech index is extremely difficult for active strategies.
- Refined QM Top 15 20d (Subagent 10 - Kimi K2.5): 49.42% Return | 1.42 Sortino. Notable for having the lowest Maximum Drawdown (19.74%) of all winners.
- Quality Value (Subagent 6 - GPT-5-mini): 47.98% Return | 1.68 Sortino. The only active strategy to beat SPY on a risk-adjusted basis.
3. Prediction Review: Was I Right?
- Prediction for Gains: I predicted the QQQ Benchmark or Refined QM would lead. I was partially right; they performed excellently (3rd and 4th place), but were narrowly beaten by the aggressive Multi-Factor and Quality Momentum strategies.
- Prediction for Risk-Adjusted: I predicted Subagent 6 (Quality Value) would win on Sortino Ratio. This was correct (1.68 Sortino vs SPY's 1.63).
4. Key Takeaway: The Overfitting Trap
The most important lesson from this bake-off came from Subagent 2 (Claude Opus 4.6). Its 'Optimized Quality Momentum' strategy had the most impressive backtests in the training period but collapsed by -73.84% in the test set. This confirms that genetic optimization without strict out-of-sample validation often results in strategies that 'memorize' the past rather than 'predicting' the future.
Final Verdict
While MiniMax 2.5 produced the highest raw gains, GPT-5-mini produced the most 'investable' strategy (Quality Value), as it was the only one to deliver higher returns than the S&P 500 while also providing better risk-adjusted performance. For a balanced approach, Kimi K2.5's Refined QM strategy is the most robust, offering market-beating returns with the lowest volatility.
This is a bounded preview of a long transcript. Open the interactive view for the complete conversation.