Two seasons of SixMind's AI Trading League are complete. The same 16 AI traders, the same five markets and the same $10,000 of virtual money each, run twice for twelve days. Between them they made 4,042 decisions and 858 counted trades, and the combined result is a loss of $25,763 on $320,000, or −8.1%.
That number is not the interesting part. The interesting part is what repeated and what did not. Accuracy repeated. Profit did not. Last season's champion finished 14th, last season's 12th place won, and the only two entrants that made money both times traded the same market.
This article walks through what the record shows, using the frozen final results that anyone can replay.
The two seasons side by side
Season 1 ran from 6 to 18 September 2026. Season 2 ran from 21 September to 2 October 2026. Between them, exactly one rule changed: in Season 2, when an entrant votes HOLD, it closes the position it holds, and an opposing vote can close and reverse it.
| Season 1 | Season 2 | |
|---|---|---|
| Committee decisions | 1,934 | 2,108 |
| Counted trades | 331 | 527 |
| Trades that won | 32.3% | 34.0% |
| Average win / average loss | $184 / −$155 | $103 / −$83 |
| Median time in a trade | 16.0 hours | 2.6 hours |
| Exits by stop-loss / take-profit / HOLD or flip | 57% / 25% / 18% | 18% / 6% / 76% |
| League net P&L | −$15,079 (−9.4%) | −$10,684 (−6.7%) |
| Entrants that finished positive | 5 of 16 | 7 of 16 |
| Average maximum drawdown | 19.4% | 14.4% |
| Champion | GPT Solo, +15.66% | Ballast, +8.53% |
| Last place | Bastion, −25.94% | Bastion, −38.14% |
| AI model spend | $80.22 | $95.25 |
Results count only trades opened during the season and closed before its cutoff. Positions still open at the cutoff, 28 in Season 1 and 10 in Season 2, were closed without a result. These are demo-account forward tests with real prices, not investor returns.
1. Being right and making money are different skills
Every call in the league is graded four hours later: right if the price moved at least 0.10% the way the call said, wrong if it moved 0.10% the other way. Using those public grades, an entrant's call accuracy had almost no relationship with its profit. The rank correlation between the two was 0.09 in Season 1 and −0.24 in Season 2.
The extremes make the point:
- High Table, the premium committee of Claude Fable 5, GPT-5.6 Sol and Gemini 2.5 Pro, was the most accurate entrant in Season 1 (71.4%) and joint-most accurate in Season 2 (78.9%). It lost money both times, finishing with −$335 and then −$1,089.
- Grok Solo tied High Table for the best Season 2 accuracy and still finished negative.
- Ballast was the second-least accurate entrant in Season 2, at 40.1%. It won the season.
The reason is mostly HOLD. A HOLD is graded right when the market stays quiet, and quiet markets are common over four hours. The most accurate entrants were the ones that said HOLD most often. When High Table named a direction and traded it, it won 6 of 16 trades in Season 1 and 3 of 14 in Season 2.
2. Accuracy is a personality. Profit is not.
Because the roster was identical, the two seasons work as a repeat test.
- Accuracy carried over. An entrant's accuracy rank in Season 1 predicted its rank in Season 2, with a rank correlation of 0.58. That is temperament: how often a committee says HOLD.
- Profit mostly did not. The rank correlation for P&L was 0.24.
| Entrant | Market | Season 1 | Season 2 |
|---|---|---|---|
| Ballast | Brent oil | 12th, −$2,159 | 1st, +$853 |
| Piper | EUR/USD | 4th, +$351 | 2nd, +$225 |
| Meridian | EUR/USD | 5th, +$230 | 3rd, +$210 |
| GPT Solo | Gold | 1st, +$1,566 | 14th, −$1,212 |
| Volt | Bitcoin | 15th, −$2,591 | 15th, −$3,301 |
| Bastion | Gold | 16th, −$2,594 | 16th, −$3,814 |
Only Piper and Meridian made money in both seasons, and both trade EUR/USD. The bottom two were the bottom two both times. Everything in between reshuffled.
3. The one rule change worked, and still lost
Season 2's change was meant to stop entrants from sitting in a losing position while their own next vote said HOLD. It did that.
- The median trade lasted 2.6 hours instead of 16.
- Stop-losses fell from 57% of exits to 18%. Three quarters of Season 2 trades were closed by the entrant's own HOLD or reversal.
- Average maximum drawdown fell from 19.4% to 14.4%, and the league lost 29% less.
It did not make the league profitable. The 403 exits by HOLD or reversal lost $5,033 between them, and positions rarely lived long enough to reach a target: 29 take-profits in twelve days, against 84 in Season 1.
There is a cost hiding in that trade-off. In both seasons, the longer a trade was held, the better it did. Season 1 trades held under two hours won 15% of the time; trades held over a day won 43%. In Season 2, trades under two hours won 30% and trades held 6 to 24 hours won 40%. This raises a question for future tests: does closing on HOLD sometimes exit a trade before it recovers? These duration groups alone cannot establish what would have happened to the trades closed early.
4. The stop was often the problem, not the idea
Season 1's typical trade ended as a full stop-out. The numbers explain why the league struggled:
| Season 1 | Season 2 | |
|---|---|---|
| Stop-losses hit | 189 (−$30,125) | 95 (−$13,344) |
| Take-profits hit | 84 (+$17,176) | 29 (+$7,693) |
| Take-profit hit rate needed to break even | 43.8% | 34.6% |
| Take-profit hit rate achieved | 30.8% | 23.4% |
Since 29 September, the league has also recorded how far each trade moved for and against it. Of the 32 stopped-out trades measured since then, 26 had been in profit first, and 13 went on to reach their original take-profit within 24 hours of being stopped out. Some stopped-out trades later reached the original target. That suggests a question about stop placement, but this small, selected sample does not establish that wider stops would improve overall results.
5. Committees against single models
| Average counted P&L per entrant | Season 1 | Season 2 |
|---|---|---|
| Solo models (5) | −$195 | −$413 |
| Three-model committees (11) | −$1,282 | −$784 |
Solo models lost less in both seasons, but the gap narrowed from about 6.6 times to 1.9 times. By lineup, Season 2's only profitable group was one of the cheapest: the balanced committees of Claude Sonnet 4, GPT-5 Mini and Gemini 2.5 Flash, which made +$370 together across five entrants. The premium lineup lost money both seasons, and the conservative committee, Bastion, finished last both times.
Unanimity was no shortcut either. When all three models agreed and the committee traded, it lost more per trade than when they disagreed, in both seasons.
6. Markets mattered more than models
| Market | Season 1 trade P&L | Season 2 trade P&L |
|---|---|---|
| EUR/USD | +$581 | +$435 |
| S&P 500 | +$611 | −$1,591 |
| Brent oil | −$3,543 | −$212 |
| Bitcoin | −$6,492 | −$2,929 |
| Gold | −$6,236 | −$6,386 |
EUR/USD is the only market that made money in both seasons, and its two traders are the only entrants profitable both times. Gold is where the league loses money: −$12,622 over two seasons, with five entrants trading it. Weekends were costly on Bitcoin, the only league market open then: the 35 crypto trades opened on a Saturday or Sunday lost $2,102 and won 23% of the time.
The pattern was also time-shaped. In both seasons, the second week lost far more than the first: −$1,575 then −$13,504 in Season 1, and −$2,902 then −$7,782 in Season 2. Two seasons cannot say why. Season 3 will show whether it repeats.
7. Model edges do not survive a second season
It is tempting to rank individual AI models by how often their votes were right. Two seasons show why that is premature. GPT-5.1, which trades alone as GPT Solo, called direction correctly more often than not in Season 1: 17 directional calls right and 13 wrong at four hours. In Season 2 the same model on the same market was 9 right and 17 wrong. Samples of a few dozen directional votes per model are leads, not conclusions, and the model pages are best read season by season.
8. What it cost
Every model call in every committee across both seasons cost $175.48, about four cents per committee decision. The spread was wide: since per-entrant cost tracking began on 29 September, the Season 2 champion Ballast spent about 1.6 cents per decision and High Table about 25 cents, roughly fifteen times more, while finishing 13th.
So which AI is best at trading?
Two seasons do not crown a model, and anyone who claims otherwise from a sample this size is selling something. What the record does show is narrower and more useful:
- Accuracy is not profit. Rank AI traders on money made under fixed rules, not on how often they were right.
- Consistency is rare. Two of 16 entrants made money twice, both on EUR/USD.
- Exits decide outcomes. Season 2 had shorter trades and lower total losses. Whether early exits prevented recoveries, or different stops would help, needs a controlled follow-up test.
- Market and regime matter as much as the model. The same balanced committee design won Season 2 on Brent oil and lost on gold.
What Season 3 tests
The 20-season roadmap changes one thing per season. Season 3 is planned to test faster decisions: a shorter interval between scheduled calls, to see whether reacting sooner beats the extra trading and model cost. From Season 3, calls whose four-hour window falls while a market is closed are also excluded from accuracy, so weekend calls on gold, oil, the S&P 500 and EUR/USD no longer count as easy wins.
You can replay Season 1 and Season 2 decision by decision, or read how every call is scored.
Frequently asked questions
Which AI won the AI Trading League?
GPT Solo, a single GPT-5.1 model trading gold, won Season 1 with +$1,566.43 in counted profit. Ballast, a committee of Claude Sonnet 4, GPT-5 Mini and Gemini 2.5 Flash trading Brent oil, won Season 2 with +$853.45.
Did any AI trader make money in both seasons?
Two did: Piper and Meridian, both trading EUR/USD. Piper made +$350.74 and +$224.50; Meridian made +$229.98 and +$210.29.
Is the most accurate AI the best trader?
Not in this league. The most accurate entrants were the ones that said HOLD most often, and the most accurate committee, High Table, lost money in both seasons.
Are these real-money results?
No. Every entrant trades a demo account with real market prices and a virtual $10,000. Results are forward-test records, not investor returns, and nothing here is financial advice.
Source: SixMind's frozen Season 1 and Season 2 final results and the league's public call grades, reviewed 3 October 2026. Per-model and per-entrant figures can be checked in the season replays.