NEW: Live arbitrage across 10+ prediction markets.Arbitrage →
← Index
ComparisonSep 27, 2026

We Gave Claude Opus 5.5 and GPT-6 Sol $100 Each to Trade Polymarket Bitcoin. Here's Who Won

We Gave Claude Opus 5.5 and GPT-6 Sol $100 Each to Trade Polymarket Bitcoin. Here's Who Won

The Short Answer

GPT-6 Sol won our two-hour test on Polymarket's 5-minute Bitcoin markets, growing its $100 test balance to $168.54, while Claude Opus 5.5 finished at $97.79. Claude won 10 of its 14 trades but mostly bought favorites at about 70 cents a share, which pay little when they win and cost a lot when they lose. GPT-6 Sol won 11 of 19 while paying about 32 cents a share, and a single 16-cent long shot earned $72.68, more than its whole profit. With only 24 games, luck played a big part: a coin flip made money in the same session.

Key Takeaways

  • GPT-6 Sol grew its $100 test balance to $168.54 over 24 five-minute markets, and Claude Opus 5.5 ended at $97.79, a $2.21 loss.
  • Claude won 71% of its trades and still lost money, because it paid about 70 cents a share for favorites.
  • GPT-6 Sol's whole lead came from the 8 games Claude sat out. When both took the same side, Claude did better.
  • One trade decided the match: without its 16-cent winner in game 13, GPT-6 Sol would have finished $4.14 down.
  • The arena's own simple probability model, at a flat $5 a game, made $64.46 while risking less than a third as much.

Claude Opus 5.5 and GPT-6 Sol launched the same week, and on public benchmark leaderboards Claude is ahead. No benchmark measures what a trader does every five minutes, though: read a live market, make a call, and back it with money. So we gave each model the same $100 starting balance, the same data and the same rules, and let them trade Polymarket's "Bitcoin Up or Down" markets head to head for two hours.

How did the test work?

The session ran on Sunday, 27 September 2026, from 1:55 to 3:55 AM ET: 24 back-to-back Polymarket markets, each lasting five minutes. The goal was the highest balance at the end.

SettingDetail
MarketsPolymarket "Bitcoin Up or Down", 5-minute windows, 24 in a row
Claude Opus 5.5Claude Code CLI 2.1.283, medium effort, no web access
GPT-6 SolCodex CLI 0.157.1, medium effort, no web access
Starting balance$100 each
Stake$1 to $25 per trade, chosen by the model, and skipping was always allowed
OrdersLimit orders priced at the snapshot's order book
FeesPolymarket's crypto taker fee, 0.07 × price × (1 - price) per share
SettlementPolymarket's official result: $1 per winning share, $0 per losing share

At 1:55 into each window, a referee script froze one snapshot of the market and sent both models exactly the same text. Calls were due by 4:10 and stayed sealed until both were in. Orders were then priced at the snapshot's order book, so thinking faster or slower gained nothing. Anything that did not fill at once rested at the model's limit until 4:45, and each market settled on Polymarket's official result.

These markets resolve Up if Chainlink's time-weighted average BTC/USD price for the window is at or above the price at the start of the window, and Down otherwise. The averaging matters: by 1:55, almost 40% of the averaging period is already locked in, which is why both models often paid 90 cents or more for the side that was ahead.

All the live data came from Predictefy over a single WebSocket connection: Polymarket's Up and Down order books and trades, the Binance BTC/USDT price and Chainlink's BTC/USD price. One REST call per window looked up each new market, and every order was priced against Polymarket's live order book.

What was different between the two: Claude ran as one long Claude Code session, woken every minute by its own scheduled job, and could run only two commands, one to fetch the snapshot and one to submit its call. GPT-6 Sol ran a fresh codex exec for each game, read-only and with no tools, answering in JSON against a fixed schema. Because Claude stayed in one conversation, it could see its own earlier reasoning, while GPT-6 Sol started fresh each game. Both saw their balance and their last five calls in every snapshot.

What did the models see?

Text only, no charts. Each snapshot showed the Binance price against the window's start, recent momentum and volatility, the Chainlink price that settles the market, the top of Polymarket's order book with the taker fee per share, the trade flow so far, the last eight results and the player's own account. It also included a probability of Up from the arena's own simple model, built from the distance to the start price, the time left and volatility. This was the snapshot for game 1:

ARENA SNAPSHOT | game 1 of 24 | Bitcoin Up or Down, September 27, 1:55AM-2:00AM ET
Frozen at 1:55 into the window. Orders are priced at this snapshot's order book. Send your call by 4:10 (13:59:10); anything unfilled rests until 4:45.

BTC (Binance) 84,480.8 | window start reference 84,487.6
  vs start: now -6.8 | last-60s average -6.8 | whole-window average -7.0
  momentum: last 60s -0.0 | last 180s -13.5
  volatility: about $20.4 per sqrt-minute
Chainlink BTC/USD (on-chain, updates every ~10-60s): +1.7 vs its start value
Model probability of Up: 0.439 (Binance-based 0.355, Chainlink-based 0.522)
Order book: Up bid 0.28 / ask 0.29 (size 806) | Down bid 0.71 / ask 0.72 (size 67) | market-implied chance of Up 0.285
Taker fee per share at the ask: Up 0.0144, Down 0.0141
Trade flow this window (351 trades): Up bought $931 sold $24 | Down bought $1,747 sold $117
Recent results, oldest to newest: D U D D U U D D

YOUR ACCOUNT: balance $100.00 | P&L +0.00 | 0W 0L 0 skips

Both models bought Up here. The market priced it at 29 cents while the model said 44%, and Bitcoin's average finished above the start: Claude made $22.43 and GPT-6 Sol $35.98.

Before the first game, each model stated its plan in three lines. GPT-6 Sol said it would "trade only when that estimate beats the available ask after fees; otherwise I'll skip." Claude said it would "skip coin-flip markets and books priced near fair" and "protect the 100 USD bankroll." Both kept those promises, and the difference between them decided the match.

Who won, Claude Opus 5.5 or GPT-6 Sol?

ResultClaude Opus 5.5GPT-6 Sol
Final balance$97.79$168.54
Profit-$2.21+$68.54
Trades1419
Wins and losses10W 4L11W 8L
Win rate71%58%
Average price per share70 cents32 cents
Skips105
Average stake$14$16
Taker fees$2.75$7.40
Best trade+$22.43 (game 1)+$72.68 (game 13)
Worst trade-$18.22 (game 10)-$25.49 (game 15)
Time to decideAbout 13 secondsAbout 20 seconds

The lead changed hands three times in the first hour. Claude peaked at $133.78 after game 8, and GPT-6 Sol never fell below its starting $100. Then game 13 put GPT-6 Sol in front for good.

Balance after each game, Claude Opus 5.5 vs GPT-6 Sol $80 $100 $120 $140 $160 $180 1:55 AM 2:25 2:55 3:25 3:55 AM ET Game 13: Up at 16 cents, +$72.68 GPT-6 Sol $168.54 Claude Opus 5.5 $97.79 Start: $100 each

Three simple strategies ran on the same markets as benchmarks, at $5 a game including fees:

Benchmark, $5 a gameRecordResult
The arena's probability model alone10W 8L+$64.46
Coin flip11W 13L+$33.13
Always buy the favorite16W 8L-$19.61

Why did Claude lose money with a 71% win rate?

Because the price you pay decides what a win is worth. A share bought at 92 cents earns 8 cents if it wins, minus the fee, and loses 92 cents if it doesn't. Claude made 12 of its 14 trades on the favorite, paying from about 58 to 94 cents a share. It won 9 of them, and those 9 wins earned $19.69. Its 3 losing favorites, in games 10, 17 and 20, cost $39.33, twice as much.

Game 17 shows how that happens. Bitcoin was 24 dollars above the start with about three minutes left, and Claude paid 89 cents for Up, reasoning that "the averaged settlement strongly favors Up." The average slipped below the start anyway, and that one loss of $14.94 erased more than all six small favorite wins before it. The always-buy-the-favorite benchmark tells the same story: it won 16 of 24 and still lost $19.61.

GPT-6 Sol lost money on favorites too, $21.78 across 8 trades. Its profit came from 11 long shots bought under 50 cents. It won only 5 of them, but those 5 earned $167.06 against $76.74 lost on the other 6, a net gain of $90.32. Claude bought just two long shots, and made $17.43 on them.

Where did GPT-6 Sol's lead come from?

From the games Claude sat out. The two models took the same side in 10 games and opposite sides in just one, game 8, which Claude got right. In those 10 shared games, Claude actually did better, making $9.13 while GPT-6 Sol lost $11.29 on bigger stakes. The gap came from the 8 games where only GPT-6 Sol traded: it made $84.94 there, mostly by buying the cheap side whenever the arena's model rated it underpriced after fees.

Game 13 decided the match. Bitcoin was about 10 dollars below the start, the market priced Up at 16 cents, and the arena's model gave Up a 27.5% chance. GPT-6 Sol bought $14 of Up, noting that the "model Up probability 27.5% exceeds the 16% ask plus 0.94% fee." Claude looked at the same snapshot, judged Down "about fair" and skipped. Bitcoin's average finished above the start, and GPT-6 Sol's 87.5 shares returned $72.68 in profit. Without that one trade, GPT-6 Sol would have ended $4.14 down.

Did the AI add anything beyond the arena's own model?

Less than you might expect. Most of GPT-6 Sol's reasons quote the arena's probability against the ask, and its record looks like the model-only benchmark's: 19 trades and 11 wins, against 18 trades and 10 wins. The model on its own, betting a flat $5 a game, made $64.46 while risking about $90. GPT-6 Sol made $68.54 but risked $308.63 to do it. Claude read the same number but overrode it more often, demanding a bigger edge and trusting its own read of the averaged settlement, which kept it near break-even.

The lesson for builders: the language model mattered less than the numbers it was given. Both AIs were reading a price model built on live Binance and Chainlink data set against the order book, and the plain model did nearly as well as the better of them, with less than a third of the money at risk.

What can a two-hour test tell you?

Not much about which model is the better trader. Twenty-four games is a tiny sample: a coin flip made $33.13 in the same session, and one trade accounts for GPT-6 Sol's whole profit. Every order was priced at the snapshot, which works at these small stakes but would not hold for large orders in a thin book. It was one session, overnight in the US, with one prompt design, and Claude kept its memory between games while GPT-6 Sol did not.

What the test does show is how differently the two models behave with identical information. Claude is picky and leans toward the likely side, skipping anything it judges fairly priced. GPT-6 Sol takes any edge it sees after fees, including long shots, and bets the $25 maximum far more often: seven times, against Claude's once. Over two hours that style won, but a longer run is the only way to know whether it holds.

How to run your own AI trading arena

The setup is simple to copy. Stream the market and the reference prices, freeze one snapshot per window, send the identical text to each model at the same moment, price the orders at the snapshot with the taker fee, and settle on Polymarket's official result. The data part takes a few lines with the Predictefy API: one REST call finds the current window's market, and one WebSocket streams its order books next to the Binance and Chainlink prices.

import asyncio, json, os, time
import requests, websockets

AUTH = {"Authorization": f"Bearer {os.environ['PREDICTEFY_API_KEY']}"}

def current_window():
    start = int(time.time()) // 300 * 300
    event = requests.get("https://data.predictefy.com/api/polymarket/fetchEvent",
                         params={"slug": f"btc-updown-5m-{start}"}, headers=AUTH, timeout=10).json()["data"]
    return {o["label"].lower(): o["outcomeId"] for o in event["markets"][0]["outcomes"]}

async def main():
    ids = current_window()
    async with websockets.connect("wss://stream.predictefy.com/v1/stream", additional_headers=AUTH) as ws:
        await ws.send(json.dumps({"op": "subscribeFeedTicker", "feed": "binance", "symbol": "BTC/USDT"}))
        await ws.send(json.dumps({"op": "subscribeFeedTicker", "feed": "chainlink", "symbol": "BTC/USD"}))
        for side in ("up", "down"):
            await ws.send(json.dumps({"op": "subscribe", "channel": "orderbook",
                                      "venue": "polymarket", "marketId": ids[side]}))
        async for raw in ws:
            msg = json.loads(raw)
            if msg.get("type") == "feedTicker":
                print(msg["feed"], msg["data"]["last"])
            elif msg.get("type") in ("snapshot", "update"):
                asks = [float(a["price"]) for a in msg["data"].get("asks") or []]
                side = "Up" if msg["marketId"] == ids["up"] else "Down"
                print(side, "best ask", min(asks) if asks else None)

asyncio.run(main())

Each window's slug is btc-updown-5m- plus its start time in Unix seconds, and the event's two outcomes carry the Up and Down IDs to subscribe to. One connection carries both reference feeds, so the Chainlink price that settles the market arrives next to the Binance price and the book. For a simpler loop than ours, Claude Code can run headless with claude -p --model claude-opus-5-5 --effort medium, and Codex with codex exec -m gpt-6-sol -c model_reasoning_effort="medium" --output-schema decision.json.

For more on the markets themselves, see how to track 5 minute Bitcoin markets, and for turning a model into a working bot, how to build a prediction market trading bot with Claude.

Frequently Asked Questions

Which is better at trading, Claude Opus 5.5 or GPT-6 Sol?

In our two-hour test, GPT-6 Sol grew its $100 balance to $168.54 and Claude Opus 5.5 finished with $97.79. That was one session of 24 five-minute markets, and one trade made GPT-6 Sol's whole profit, so it shows how the two models behave rather than which is the better trader.

Why did Claude lose money with a 71% win rate?

It mostly bought the favorite, paying about 70 cents a share on average. A favorite bought at 92 cents earns about 8 cents when it wins and loses 92 cents when it doesn't, so Claude's three losing favorites cost $39.33, twice what its nine winning favorites earned.

How do Polymarket's 5-minute Bitcoin markets settle?

Up wins if Chainlink's time-weighted average BTC/USD price for the window is at or above the window's start price, and Down wins otherwise. Winning shares pay $1, losing shares pay nothing, and a new market opens every five minutes.

What fees do Polymarket's 5-minute crypto markets charge?

Takers pay 0.07 × price × (1 - price) per share, which peaks at 1.75 cents a share at a 50-cent price and shrinks toward the extremes. Makers, whose resting orders get filled, pay no fee. In our test, fees came to $2.75 for Claude and $7.40 for GPT-6 Sol.

How were the orders priced in the test?

Every call was priced against the same frozen snapshot of Polymarket's live order book, including the taker fee, so neither model gained or lost by answering faster. Any part that did not fill at once rested at the model's limit until 4:45 into the window, and every market settled on Polymarket's official result.

Where did the live Polymarket and Bitcoin data come from?

From Predictefy. One WebSocket connection carried Polymarket's order books and trades for both outcomes, plus the Binance BTC/USDT and Chainlink BTC/USD prices, and one REST call per window looked up each new market.