Arena secures $200M Series B at a $3.1B valuation

Arena, the AI model-ranking platform that began as a UC Berkeley research project in 2023, has raised a $200 million Series B at a $3.1 billion valuation. Lightspeed Venture Partners and Khosla Ventures led the round, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis and other investors participating.
The financing follows Arena’s statement that it reached $100 million in annualized run-rate revenue in June. In January, the company announced a $150 million Series A at a $1.7 billion post-money valuation and said annualized revenue then stood at $30 million. The latest round therefore brings its valuation close to twice the January level in roughly 10 months.
Community feedback becomes an evaluation product
Arena operates a free consumer platform where people submit prompts or request vibe-coded projects, compare model outputs and select the model they believe performed better. The company says the service attracts tens of millions of monthly visitors.
In September last year, Arena introduced AI Evaluations, its commercial offering for model labs and enterprises. The service supplies detailed performance analytics built from feedback gathered through its community. That model gives organizations a way to examine model behaviour beyond a single standardised test result.
The company said AI is progressing faster than the ability to evaluate it, while static benchmarks can lose value once models recognise that they are being tested. It also said enterprises increasingly need help identifying the model that fits their own internal requirements rather than relying solely on general-purpose benchmark scores.
Alignment joins the leaderboard
Arena has added alignment as a new leaderboard category. Its assessments include unauthorised action, in which a model takes action it was not asked to take; false attribution, where statements or facts are credited to the wrong source; and deceptive completion, where a model claims a task was completed when it was not.
The preliminary alignment leaderboard currently has a group of OpenAI models at the top. Claude Opus 5.5 ranks sixth and Claude Fable ninth. The category places behavioural and safety-related questions alongside the capability comparisons for which Arena is known.
What the funding signals for buyers
The Series B reflects demand for evaluation methods that use real interactions as well as static testing. For businesses selecting AI systems, benchmark rankings can provide a starting point, but AI Evaluations and the alignment measures underline the need to test candidates against the organisation’s own workflows, users and failure modes before deployment.

