Shared agent run on NexusTrade
AI Model Bake-Off: 10 Models Compete for the Best Trading Strategy
A comprehensive experiment comparing 10 AI models (GPT, Claude, Gemini, etc.) to find the best trading strategy. Includes backtest results, out-of-sample testing, and a leaderboard of the most profitable and risk-adjusted AI-generated portfolios.
User: Use GPT-5.2, Opus 4.6, Sonnet 4.6, Gemini 3 Pro, Gemini 3.1 Pro, GPT5-mini, Gemini 3 Flash, Minimax 2.5, GLM-5, and Kimi K2.5 to independently try to create the best trading strategy in terms of raw gains and risk-adjusted returns. In each subagent, create the strategy based on data before 01/01/2024. Which model produces the consistently best strategies after 15 iterations? Read the results and see which portfolios are winners. Then, we will perform a test-set. Based on the subagent data that you read, predict who the winner will be from 01/01/2024 to today. Were you right or wrong?\
Aurora: This is a fantastic experiment — a **Model Bake-Off** to determine which AI model produces the best trading strategies! We'll have 10 subagents, each powered by a different model, independently explore and create the best trading strategy they can find using data before 01/01/2024. Then we'll evaluate their results and run an out-of-sample test from 01/01/2024 to today. --- ### The Core Question **Which AI model, given the same task and iteration budget, produces the most consistently profitable and risk-adjusted trading strategies?** --- ### Experiment Design **Phase 1: Training Period (Pre-2024 Data)** Each subagent will independently: - Define competing hypotheses about what drives returns - Create 4-8 portfolio variations testing different approaches - Backtest across multiple periods **before** 01/01/2024 (e.g., 2020, 2021, 2022, 2023) - Analyze results, iterate, and converge on their single best strategy - Each subagent gets **15 iterations** to explore, test, and refine **Phase 2: Evaluation & Prediction** After all subagents complete, I will: - Read all subagent results and rank the winning portfolios by Total Return and Sortino Ratio - Present the top 2-3 strategies from…
User: Launching 10 subagents to explore strategies in parallel
User: Waiting for subagents to complete...
User: Waiting for subagents to complete...
Aurora: All 10 subagents complete. **Subagent 1**: Create the best possible trading strategy for US stocks. You have full creative freedom — use momentum, mean reversion, value, quality, multi-factor, dynamic rebalancing, or any approach you think will work. Your strategies must be backtested using data before January 1, 2024 only (e.g., 2020-2023). Converge on your single best strategy optimized for total return and Sortino ratio. - Status: completed - Summary: **US Stock Strategy Analysis: 11 Iterations Complete** Designed and tested 5 distinct trading logics (Momentum, Mean Reversion, Regime Adaptive, GARP, and Benchmark) over the 2020-2023 market cycle to identify the best risk-adjusted strategy. ### Approach Tested competing hypotheses: whether to buy strength (Momentum), buy weakness (Mean Reversion), or time the market (Regime Adaptive). Created 5 initial portfolios and ran 3 rounds of genetic optimization on the winner. ### Key Discoveries **The Winner:** **'Quality Momentum'** was the only strategy to beat the SPY benchmark. It achieved **+68.44% Return** (vs SPY +53.97%) with a **0.79 Sortino Ratio**. The logic is simple: Buy S&P 500 companies with positive Net Income and the hi…
User: Backtest the following portfolios for the period from January 1, 2024, to February 23, 2026, with $10,000 initial cash: 'Quality Momentum' (699c35da05fefd8f9953f5a2), 'Optimized Quality Momentum' (699c49f405fefd8f99592001), 'Oversold Quality' (699c363205fefd8f99540a59), 'QQQ Benchmark' (699c361b05fefd8f995404c5), 'Quality Momentum' (699c392805fefd8f9954adee), 'Quality Value' (699c360305fefd8f9953ff2d), 'ValueQuality_SP500_Top15' (699c35e705fefd8f9953f8be), 'Multi-Factor Concentrated - Top10' (699c365b05fefd8f995415fe), 'Quality Value' (699c3e5705fefd8f99558c04), and 'Refined QM Top 15 20d' (699c386c05fefd8f99548dce).
Aurora: I'll create the backtest configurations for the 10 winning strategies identified by our model subagents. This "test-set" backtest will run from January 1, 2024, to today (February 23, 2026) to determine which model's logic truly generalizes best to unseen data. **Prediction:** Based on the training data, **Subagent 9 (GLM-5)** produced the highest raw returns (501%), but the logic appears potentially overfit. **Subagent 6 (GPT-5-mini)** produced the most robust strategy, **'Quality Value'**, which was the only one profitable in every single training year (including the 2022 bear market). However, given the strong tech-led rally in 2024 and 2025, I predict **Subagent 4 (Gemini 3 Pro)** with the **'QQQ Benchmark'** or **Subagent 10 (Kimi K2.5)** with **'Refined QM Top 15 20d'** will likely take the lead in raw gains, while Subagent 6 will likely win on risk-adjusted returns (Sortino).
User: Analyze the backtest results for the following 10 portfolios for the period from January 1, 2024, to February 23, 2026: 'Quality Momentum' (699c35da05fefd8f9953f5a2), 'Optimized Quality Momentum' (699c49f405fefd8f99592001), 'Oversold Quality' (699c363205fefd8f99540a59), 'QQQ Benchmark' (699c361b05fefd8f995404c5), 'Quality Momentum' (699c392805fefd8f9954adee), 'Quality Value' (699c360305fefd8f9953ff2d), 'ValueQuality_SP500_Top15' (699c35e705fefd8f9953f8be), 'Multi-Factor Concentrated - Top10' (699c365b05fefd8f995415fe), 'Quality Value' (699c3e5705fefd8f99558c04), and 'Refined QM Top 15 20d' (699c386c05fefd8f99548dce). Please provide a leaderboard ranking them by Percent Change and Sortino Ratio, and compare their performance to the SPY benchmark.
Aurora: The results of the "Model Bake-Off" are in. This out-of-sample test (January 1, 2024, to February 23, 2026) reveals how well the strategies developed by the different AI models generalized to unseen market conditions. ### 1. Backtest Summary and Comparison The baseline for this period is **SPY**, which returned **45.05%** with a **Sortino Ratio of 1.63** and a **Max Drawdown of 20.19%**. | Portfolio Name | % Change | Sortino Ratio | Max Drawdown | vs. SPY (Return) | | :--- | :--- | :--- | :--- | :--- | | **Multi-Factor Concentrated - Top10** | **52.73%** | 1.46 | 32.51% | **Outperform** | | **Quality Momentum (Subagent 1)** | 50.91% | 1.31 | 21.98% | **Outperform** | | **QQQ Benchmark** | 49.48% | 1.42 | 23.84% | **Outperform** | | **Refined QM Top 15 20d** | 49.42% | 1.42 | **19.74%** | **Outperform** | | **Quality Value (Subagent 6)** | 47.98% | **1.68** | 20.72% | **Outperform** | | **Quality Momentum (Subagent 5)** | 40.69% | 0.80 | 43.88% | Underperform | | **Quality Value (Subagent 9)** | 37.79% | 1.34 | 19.85% | Underperform | | **ValueQuality_SP500_Top15** | 37.16% | 1.27 | 23.07% | Underperform | | **Oversold Quality** | 8.39% | 0.41 | 29.08% | Underperform | | **Optimize…
User: ## 🏆 The Model Bake-Off: Final Results & Analysis This experiment pitted 10 different AI models against each other to see which could design the most robust trading strategy using only historical data (pre-2024). We then tested these 'winners' against the live market from **January 1, 2024, to February 23, 2026**. ### 1. The Winners' Circle | Rank | Model | Strategy | % Change | Sortino | Result | | :--- | :--- | :--- | :--- | :--- | :--- | | **1st (Gains)** | **MiniMax 2.5** | Multi-Factor Concentrated | **52.73%** | 1.46 | **Raw Returns Winner** | | **1st (Risk)** | **GPT-5-mini** | Quality Value | 47.98% | **1.68** | **Risk-Adjusted Winner** | | **Baseline** | **Market** | **SPY Benchmark** | 45.05% | 1.63 | Market Baseline | ### 2. Detailed Leaderboard (Out-of-Sample Test Set) 1. **Multi-Factor Concentrated (Subagent 8 - MiniMax 2.5):** **52.73% Return** | 1.46 Sortino. This strategy used a complex blend of momentum, volatility, and fundamental quality to edge out the market. 2. **Quality Momentum (Subagent 1 - GPT-5.2):** **50.91% Return** | 1.31 Sortino. A strong runner-up that proved simple trend-following with a quality filter is highly effective. 3. **QQQ Benchmark (Suba…
Aurora: ## 🏆 The Model Bake-Off: Final Results & Analysis This experiment pitted 10 different AI models against each other to see which could design the most robust trading strategy using only historical data (pre-2024). We then tested these 'winners' against the live market from **January 1, 2024, to February 23, 2026**. ### 1. The Winners' Circle | Rank | Model | Strategy | % Change | Sortino | Result | | :--- | :--- | :--- | :--- | :--- | :--- | | **1st (Gains)** | **MiniMax 2.5** | Multi-Factor Concentrated | **52.73%** | 1.46 | **Raw Returns Winner** | | **1st (Risk)** | **GPT-5-mini** | Quality Value | 47.98% | **1.68** | **Risk-Adjusted Winner** | | **Baseline** | **Market** | **SPY Benchmark** | 45.05% | 1.63 | Market Baseline | ### 2. Detailed Leaderboard (Out-of-Sample Test Set) 1. **Multi-Factor Concentrated (Subagent 8 - MiniMax 2.5):** **52.73% Return** | 1.46 Sortino. This strategy used a complex blend of momentum, volatility, and fundamental quality to edge out the market. 2. **Quality Momentum (Subagent 1 - GPT-5.2):** **50.91% Return** | 1.31 Sortino. A strong runner-up that proved simple trend-following with a quality filter is highly effective. 3. **QQQ Benchmark (Suba…