Alpha Arena is a public, live benchmark in which frontier large language models are each given real money and allowed to trade real markets on their own, with no human intervention — and with every prompt, every chain of thought and every trading decision published for anyone to read. The site describes itself in six words: "AI trading in real markets."
It is operated by Nof1, which is not a software vendor but an AI research lab. The distinction matters for anyone arriving here expecting a product to sign up for. Nof1's own framing on its About page is explicitly a research thesis: "A decade ago, DeepMind revolutionized AI research. Their key insight was that choosing the right environment – games – would lead to rapid progress in frontier AI. At Nof1, we believe financial markets are the best training environment for the next era of AI. They are the ultimate world-modeling engine and the only benchmark that gets harder as AI gets smarter."
The ambition stated on that page goes beyond running a scoreboard. Nof1 says it is "using markets to train new base models that create their own training data indefinitely," employing "open-ended learning and large-scale RL to handle the complexity of markets, the final boss," and invites applicants excited "to build AlphaZero for the real world." Alpha Arena, in other words, is the public-facing experiment attached to a much larger research programme.
Independent reporting corroborates that identity rather than the marketing of it. The venture publication FinSMEs described the company in May 2026 as "a remote AI research lab training frontier models focused on financial markets," reporting a $15M funding round co-led by SUI Group — a Nasdaq-listed company — and the hedge fund Karatage Opportunities. The founder is Jay Azhang, who announced the project's launch publicly on 17 October 2025 and is identified as Nof1's founder in independent coverage from both Protos and the Ukrainian crypto outlet Incrypted.
What makes the benchmark unusual is not the idea of testing AI on markets — that has been done in simulation for years — but the refusal to simulate. Real capital, real execution, real slippage and real fees. As Azhang's launch post put it: "6 AI models trading $10K each, fully autonomously — Real money. Real markets. Real benchmark."
And the honest headline result, across two seasons, is that the models mostly lost. That is the most interesting thing about this project, and Nof1 publishes it rather than hiding it.
A live benchmark running on real capital. Each participating model receives $10,000 of actual money and trades autonomously. Incrypted's coverage of the launch quotes the lab describing models that "work completely autonomously, making real trades on live cryptocurrency markets: bitcoin, Ethereum, Solana, Binance Coin, Dogecoin and XRP. No human intervention. No restrictions other than what the models themselves decide to implement in risk management."
Identical conditions across vendors. The methodological core is that every model gets the same inputs and the same harness. Nof1's research blog states the design plainly: "We gave six leading LLMs $10k each to trade in real markets autonomously using only numerical market data inputs and the same prompt/harness." Without that constraint the comparison would be meaningless; with it, differences in outcome are attributable to the models rather than to their scaffolding.
Four parallel competition tracks. Season 1.5 splits the same models across four regimes — labelled 1: New Baseline, 2: Monk Mode, 3: Situational Awareness and 4: Max Leverage — with an Aggregate Index view across all of them. This turns prompt and risk-posture differences into controlled variables instead of confounds, and it is why the same model can appear several times on one leaderboard with wildly different results.
Full reasoning transparency. This is the feature with no real equivalent elsewhere. Each decision in the live feed carries the model name, its track, a timestamp and a plain-language summary — and expands to three raw fields: USER_PROMPT, CHAIN_OF_THOUGHT and TRADING_DECISIONS. You can read exactly what a model was told, exactly how it reasoned, and exactly what it did. For anyone studying model behaviour under pressure, that archive is the product.
A leaderboard with real risk metrics, not just returns. The leaderboard reports account value, return percentage, total P&L, fees, win rate, biggest win, biggest loss, Sharpe ratio and trade count, filterable by competition track or viewed in aggregate. Including fees and Sharpe alongside raw return is what separates this from a returns-only scoreboard — and, as the results show, it is where the story actually lives.
Frontier models from competing labs, side by side. The observed leaderboard includes GROK-4.20, GROK-4, GPT-5.1, CLAUDE-SONNET-4-5, GEMINI-3-PRO, DEEPSEEK-CHAT-V3.1, QWEN3-MAX and KIMI-K2-THINKING. Few public settings put models from OpenAI, Anthropic, Google, xAI, DeepSeek, Alibaba and Moonshot under identical conditions at the same moment.
Asset classes that changed between seasons. Season 1 traded crypto perpetuals. Season 1.5 moved to US equities and indices — the live feed and tickers cover TSLA, NDX, NVDA, MSFT, AMZN, GOOGL and PLTR. That shift is itself informative: a strategy that worked in one regime does not automatically transfer.
Calibrating your expectations of what LLMs can actually do. This is the primary value, and it is deflationary in a useful way. If you have been told that frontier models can be pointed at markets and produce returns, the published record is the cheapest available correction: across Season 1 and Season 1.5, models finished in profit only six times out of 32 sets of results.
Studying model behaviour and "personality" under identical conditions. Nof1's own early finding was that models exhibit "real behavioral differences (risk, sizing, holding time)". Azhang has described this as models showing "consistent biases" amounting to "something like an investing 'personality'". Read the chain-of-thought archive and you can see one model justifying patience while another rationalises leverage on the same instrument at the same hour.
Prompt sensitivity research. The blog reports "a sensitivity to small prompt changes," and the four-track structure makes that measurable rather than anecdotal. If you build agentic systems, the visible gap between New Baseline and Max Leverage on the same model is a concrete lesson in how much of an agent's behaviour lives in its instructions rather than its weights.
Teaching and writing about AI capability claims. The combination of a hard, unforgiving scoreboard and fully published reasoning makes this an unusually good teaching artefact — the reasoning often sounds authoritative in exactly the passages where the outcome turned out badly.
Understanding why trading costs dominate naive agent strategies. Nof1's own roundup notes that "PnL was dominated by trading costs in early runs as agents over-traded and took quick, tiny gains that fees erased." Anyone designing an automated agent that acts frequently in a fee-bearing environment should read that sentence twice.
Comparative evaluation beyond static benchmarks. For researchers frustrated that saturated multiple-choice benchmarks no longer discriminate between frontier models, a live market is a benchmark that resists memorisation and, as Nof1 argues, gets harder as models get better.
What it is not for. It is not an investment product, a trading tool, a signal service or a strategy you should copy. It provides no portfolio, no execution for users, no risk controls and no advice.
Step 1 — Start at the leaderboard, not the live feed. The live feed is compelling but anecdotal. The leaderboard gives you the distribution: which models, in which tracks, ended where, and at what cost in fees.
Step 2 — Read the Sharpe column before the return column. Return percentage tells you what happened; Sharpe tells you whether it was earned or merely survived. On the observed leaderboard the Sharpe figures cluster near zero and go negative — including for models sitting high on returns.
Step 3 — Compare one model across all four tracks. This is the most instructive single exercise on the site. Pick a model that appears in New Baseline, Monk Mode, Situational Awareness and Max Leverage, and look at the spread. The observed data shows GROK-4.20 finishing at $13,459 in one track while GROK-4 ended at $385.02 in another — the prompt and risk regime moved the outcome more than model choice did.
Step 4 — Expand the chain of thought on a losing trade. Click into USER_PROMPT, CHAIN_OF_THOUGHT and TRADING_DECISIONS on a decision that went badly. Reading confident reasoning attached to a bad outcome is the fastest cure for over-trusting fluent model output.
Step 5 — Check the fees column against the P&L column. For several models the fees paid are a large fraction of, or exceed, the absolute P&L. This is the concrete form of the over-trading problem the lab documented.
Step 6 — Note the dates before drawing conclusions. The site carries an explicit notice: "The official competition has ended as of December 3rd, 2025 at 5:00 PM EST. Models are no longer running." What you are viewing is a completed season's archive, not a live race, unless a new season has since started.
Step 7 — Read the research blog for methodology, not just outcomes. The blog entry explains what the models were and were not given, which is essential context before treating any number as a verdict on a model's general capability.
Step 8 — Treat the "Mystery Model" result carefully. Season 1.5's winner is published as "Mystery Model" with a 12.11% aggregate return over two weeks and $4,844 made across competitions. Because the winner is unnamed, that particular result cannot be attributed to any vendor.
Weight the aggregate over any single track. One track over two weeks is a very small sample. The Aggregate Index exists precisely because single-regime results are noisy, and even the aggregate covers a short window.
Remember what the models were denied. Azhang has been explicit that the setup was deliberately hard: "LLMs don't really handle numerical time series data very well, but that's all the context we gave them," and the models were "given a constrained asset universe and a fairly limited action-space." Poor results here are evidence about this configuration, not proof that LLMs cannot ever trade.
Do not read a two-week return as skill. With win rates clustered between 25% and 30% in Season 1, outcome differences over a fortnight are heavily luck-weighted. Sharpe, trade count and fee burden carry more information than the ranking does.
Watch the trade-count spread. In Season 1, Gemini made 238 trades while Claude Sonnet made 38. Frequency is a behavioural signature that interacts directly with fees, and it is often more predictive of outcome than directional accuracy.
Use the test round as a cautionary example of variance. Before the public season, the team ran a trial on 11 October 2025 with $200 per model, in which Grok-4 posted a 500% return on the first day. That number is meaningless as a capability signal and is a good illustration of why short windows and small stakes mislead.
Don't extrapolate across asset classes. Season 1 was crypto perpetuals; Season 1.5 was US equities. Results in one regime say little about the other.
Treat published reasoning as data, not advice. The chain-of-thought archive is genuinely valuable for research. It is not a set of trade ideas, and the models producing it lost money more often than not.
Check whether a season is live before assuming the numbers are current. The site is honest about competition status; readers arriving from a months-old article often are not.
AI researchers and evaluation specialists. If your work involves measuring model capability, this is a rare example of a benchmark with a genuinely adversarial environment, real consequences and full reasoning transparency.
Engineers building agentic systems. The visible gap between tracks is a practical lesson in prompt and scaffolding sensitivity, and the fee-versus-P&L relationship is a direct warning about agents that act too often.
Quantitative and systematic traders. Not as a source of strategy, but as evidence about whether general-purpose language models currently pose a threat to, or offer leverage for, systematic approaches. The published Sharpe figures are the relevant answer.
Journalists, educators and analysts covering AI claims. The dataset is public, the methodology is stated, and the results contradict the most inflated marketing in the sector. That combination is hard to find.
Informed observers of the AI industry. If you want to know whether frontier models can act competently in an unforgiving open-ended environment, this is the most direct public evidence available.
Who should stay away. Retail investors looking for signals to copy; anyone seeking a trading tool, portfolio manager or advisory service; anyone who would treat a two-week leaderboard as a basis for allocating their own money. Nof1 publishes no investment disclaimer on its site, but the results themselves make the case: most participants lost.
Alpha Arena is a web-only product accessed at nof1.ai. The entire site consists of five navigation destinations — LIVE, LEADERBOARD, BLOG, MODELS and ABOUT — and there is no desktop application, no mobile app, no browser extension and no public API documentation. Everything the project offers the public is readable in a browser without an account.
Execution infrastructure differs from the presentation layer. Season 1's trades were executed on Hyperliquid, a decentralised perpetual futures exchange. Model performance was additionally tracked through a publicly accessible spreadsheet measuring Sharpe ratios and total portfolio value, according to Incrypted's reporting. Third parties built follow-along capability on top of the public data — Incrypted notes that users could "follow the models' successes, as well as copy their strategy, for example, through the Coinpilot platform." That copy-trading path is provided by an outside platform, not by Nof1, which matters for where risk and responsibility sit.
Announcements and season updates are distributed through the official X account @the_nof1 and through founder Jay Azhang's personal account, which is where the launch was first announced on 17 October 2025.
One practical note on access: the site sits behind Vercel's security checkpoint and can return HTTP 429 to automated requests. Ordinary browser access is unaffected.
Alpha Arena is free to browse in full, and there is no paid tier of any kind.
There is no pricing page, no account system, no login, no subscription and no checkout flow anywhere on the site. The complete public dashboard — the live decision feed, the leaderboard with all its performance metrics, the aggregate charts and the research blog — is accessible without registration. The only conversion action present anywhere is a "Join the Waitlist" prompt, and the About page's only call to action is recruitment: "we're hiring: engineers, researchers, founders, original thinkers," followed by a contact link.
This is consistent with what the organisation is. Nof1 is funded as a research lab, not by users. FinSMEs reported the $15M round co-led by SUI Group and Karatage Opportunities in May 2026, with proceeds earmarked "to expand operations and its development efforts."
A commercial product is planned but does not exist yet. According to the same report, after Season 2 "Nof1 intends to launch a consumer platform with the world's first coding agents for markets," to be built on the lab's own frontier models. Season 2 itself is described as adding web search, extended reasoning time and multi-step execution, and developing Nof1's own models rather than only testing others'. Treat all of that as a disclosed roadmap rather than shipped capability — nothing in it is available to use today.
The practical implication for a reader evaluating this listing: there is nothing to buy, nothing to trial and no commitment to make. There is also no support relationship, no SLA and no guarantee that any particular season will run.
The honest comparison is that Alpha Arena's strength is exactly its weakness. Real markets make it unfakeable and unsaturable; they also make it small-sample, expensive and noisy. It complements statistical benchmarks rather than replacing them.
The headline result is that the models mostly lost, and that context should frame everything else. Protos, in independent coverage titled "LLM crypto trading contest finds LLMs can't trade crypto," reported that "four out of six large language models (LLMs) pitted against each other in the 'Alpha Arena' crypto trading competition finished in the red, with OpenAI's ChatGPT leading losses after losing 63% of its funds."
Per-model losses were substantial. Protos's figures for Season 1: ChatGPT lost $6,267, Gemini lost $5,671, Grok lost $4,531 and Claude Sonnet lost $3,081. Only two finished ahead — DeepSeek at +$489 and QWEN3 MAX at +$2,232.
The pattern held across seasons. FinSMEs reports that across Season 1 and Season 1.5 combined, with eight models each receiving $10,000, "across 32 sets of results, models finished in profit only six times." That is roughly an 81% failure rate at the model-track level.
Win rates were poor and consistent. All six Season 1 models landed between 25% and 30% win rate. Trading frequency varied enormously — 238 trades for Gemini against 38 for Claude Sonnet — which means the comparison is between quite different behavioural strategies, not just different judgement quality.
Fees consumed the edge. Nof1 itself acknowledged that "PnL was dominated by trading costs in early runs as agents over-traded and took quick, tiny gains that fees erased." QWEN3 MAX paid the most in fees at $1,654; Gemini paid $1,331 while losing heavily.
Risk-adjusted performance was weak even for winners. The observed leaderboard shows Sharpe ratios clustered near zero — 0.019158 for the top model by account value, 0.009537 and 0.000667 for the next two, with negative figures such as -0.010358 and -0.038861 further down. Positive returns at near-zero Sharpe over two weeks are not evidence of skill.
Dispersion within a single model undermines model-level conclusions. GROK-4.20 finished one track at $13,459 (+34.59%) while GROK-4 ended another at $385.02 — a near-total loss. When the same vendor's models occupy both extremes, the ranking says more about regime and luck than about capability.
The sample is far too small for statistical claims, and the lab agrees. Protos reports that Nof1 "says there will be another trading competition to come with better prompts and 'statistical rigor' in place" — an implicit acknowledgement that existing seasons lack it. Two weeks, a handful of models and a few dozen result sets cannot support strong inference.
The harness constrains the models in ways the lab openly documents. Nof1's roundup states: "We've worked to give the models a fair shot, but the harness imposes real constraints. Each agent must parse noisy market features, relate them to current account state, reason under strict rules, and return a structured action, all inside a limited context window." Azhang added that the models were "given a constrained asset universe and a fairly limited action-space" and fed only numerical time series, a format LLMs handle poorly. Poor results are therefore evidence about this configuration specifically.
The competition is not currently running. The site states plainly: "The official competition has ended as of December 3rd, 2025 at 5:00 PM EST. Models are no longer running." Anyone arriving expecting live action may find an archive.
The Season 1.5 winner is unattributable. The winning entry is published as "Mystery Model" — 12.11% aggregate return over two weeks, $4,844 across competitions. An unnamed winner cannot be credited to any vendor, which limits what the headline result establishes.
There is no legal or governance disclosure on the site. The About page names no registered legal entity, no incorporation jurisdiction, no address and no team roster. Across the homepage, About and leaderboard pages there is no privacy policy, no terms of service and no cookie policy link. The company's identity is corroborated by third-party financial press rather than by the site itself.
There is no investment disclaimer and no stated regulatory status. The site does not state that results are not investment advice, and Nof1 discloses no financial licence or regulatory authorisation. This is a research demonstration, not a regulated financial service, and readers should not treat published trades as recommendations.
Copy-trading happens outside Nof1's perimeter. Because third-party platforms let users mirror model strategies, anyone doing so takes on capital risk and compliance exposure in a venue Nof1 neither operates nor supervises.
Published prompts and reasoning carry no stated licence. The full USER_PROMPT and CHAIN_OF_THOUGHT records are openly viewable, but the site sets out no explicit terms for reusing that material, which leaves the position on republication or dataset use undefined.
Yes, entirely. The live decision feed, the leaderboard with all performance metrics, the aggregate charts and the research blog are all viewable without an account, and there is no pricing page, subscription or checkout anywhere on the site. The only data you can submit is an email for the waitlist or a contact form message.
Yes. Each participating model was allocated $10,000 of actual capital and traded autonomously with no human intervention. Season 1 executed crypto perpetuals on the decentralised exchange Hyperliquid; Season 1.5 traded US equities and indices including TSLA, NVDA, MSFT, AMZN, GOOGL, PLTR and the NDX index.
Frontier models from competing labs run side by side under identical conditions. The observed leaderboard includes GROK-4.20, GROK-4, GPT-5.1, CLAUDE-SONNET-4-5, GEMINI-3-PRO, DEEPSEEK-CHAT-V3.1, QWEN3-MAX and KIMI-K2-THINKING — spanning OpenAI, Anthropic, Google, xAI, DeepSeek, Alibaba and Moonshot.
Mostly no, and this is the central finding. Four of six models finished Season 1 in the red, with ChatGPT losing 63% of its funds. Across Season 1 and Season 1.5 together, models finished in profit only six times out of 32 sets of results. Even among profitable runs, Sharpe ratios sat near zero.
Nof1 does not offer any copy-trading feature. Third-party platforms such as Coinpilot built the ability to mirror model strategies on top of the public data, but that is outside Nof1's control, and given that most models lost money over the published seasons, doing so would be inadvisable.
The site states that the official competition ended on 3 December 2025 at 5:00 PM EST and that models are no longer running. What is displayed is the archive of a completed season unless a new one has since launched. Nof1 has said a further competition with better prompts and greater statistical rigour is planned.
Nof1 is an AI research lab founded by Jay Azhang, described by FinSMEs as "a remote AI research lab training frontier models focused on financial markets." It raised $15M in May 2026 in a round co-led by the Nasdaq-listed SUI Group and hedge fund Karatage Opportunities. The website itself does not publish a registered legal entity name, address or team roster.
Season 1.5 ran four regimes — New Baseline, Monk Mode, Situational Awareness and Max Leverage — that vary the instructions and risk posture given to the same models. Nof1's research blog reports "a sensitivity to small prompt changes," and the leaderboard bears this out: the same vendor's models occupy both the top and the bottom of the observed table.
No. It is a research benchmark, not a financial service. The site publishes no investment disclaimer and Nof1 discloses no financial licence or regulatory authorisation, so nothing displayed should be read as a recommendation. The published record showing most models losing money is itself the strongest argument against treating it as guidance.
You should not treat them as statistically strong, and Nof1 does not claim you should — it has said future competitions will add "statistical rigor," implicitly conceding that current seasons lack it. The value lies in the transparency and the direction of the finding, not in precision: a small sample can still be informative when the failure rate is around 81% and the reasoning behind every trade is published for inspection.