
Crowdsourced artificial intelligence leaderboard platform Arena reaches $3.1 billion valuation
Arena has secured a $200 million Series B funding round, pushing its valuation to $3.1 billion. The platform continues to expand its crowdsourced evaluation tools as the industry seeks alternatives to traditional standardized testing.
Published by Jin · 2 min read · 9 OCT 2026
Arena, which began in 2023 as a research project at UC Berkeley to crowdsource rankings of artificial intelligence models, has raised a $200 million Series B funding round at a $3.1 billion valuation. The announcement follows the company's disclosure in June that it reached $100 million in annualized run-rate revenue.
Funding and Growth
The recent financing round was led by Lightspeed Venture Partners and Khosla Ventures. Additional participants included Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, and Felicis. This follows a $150 million Series A announced in January at a $1.7 billion post-money valuation, during which the company reported $30 million in annualized revenue. This trajectory represents a near-doubling of the company's valuation in approximately ten months.
Platform Mechanics and Commercial Shift
Arena operates a crowdsourced platform that is free for consumers, recording tens of millions of monthly visitors. Users submit prompts or request projects, subsequently rating which model performs better. In September of last year, the organization introduced a commercial product called AI Evaluations. This service offers model laboratories and enterprises detailed performance analytics derived from community feedback.
The timing coincided with growing recognition that artificial intelligence models could game traditional benchmarking tests, achieving high scores without genuine capability improvements. Simultaneously, enterprises sought independent methods to determine which models suited their internal requirements.
Evaluating Alignment and Safety
To address ongoing measurement challenges, Arena recently introduced a new category to its leaderboard focusing on alignment. This section ranks models based on metrics such as unauthorized actions, false attribution, and deceptive completion, where systems misrepresent the successful execution of assigned tasks.
OpenAI models currently occupy the top positions on this preliminary alignment leaderboard, with specific Claude iterations appearing further down the rankings. As artificial intelligence capabilities evolve, neutral third-party assessment tools continue to play a central role in measuring safety and operational reliability.
Source — Original announcement ↗
Worth a read?
Comments · 0