Global Startups Insights
Lets your startup
California, Funding News, USA Startup Funding News

Arena Raises $200 Million Series B at $3.1 Billion Valuation

October 9, 2026
By
Loren Baker
Arena Raises $200 Million Series B at $3.1 Billion Valuation

Arena has raised $200 million in Series B funding at a $3.1 billion valuation to expand its AI evaluation platform and help businesses and researchers understand how well AI models and agents perform in real-world situations.

The San Francisco-based company announced that the round was co-led by Lightspeed Venture Partners and Khosla Ventures. Other participants included Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, Andreessen Horowitz (a16z), Felicis, AMP PBC, QuantumLight and The House Fund.

Arena plans to use the funding to improve its real-world AI evaluation capabilities, expand its platform and develop its new Alignment Index, which measures whether AI systems behave in ways that align with human instructions and expectations.

The funding follows a period of rapid growth for the company. Arena says it has exceeded $100 million in annualized revenue since raising its Series A in January 2026.

Arena’s Growth in AI Evaluation

Arena is building an independent platform to evaluate AI models and agents using real-world interactions. Its platform allows users to compare AI systems across different tasks and capabilities, helping researchers and the wider AI community understand how models perform beyond traditional benchmarks.

According to the company, its growth since the Series A includes:

  • $100 million-plus in annualized revenue, according to Arena.
  • 7 million sessions on Agent Arena in less than five months after launch.
  • 350 million sessions across the wider Arena platform.
  • 62 million votes across text, vision, code, search, video and image categories.
  • More than 1,000 model evaluations covering open and proprietary AI models.
  • 375,000 open-sourced data points to support AI research and analysis.

Arena also reports tens of millions of monthly visitors from more than 150 countries.

These figures reflect the scale of the company’s evaluation platform and its focus on collecting feedback and performance data from real users. The company aims to help people compare AI systems on practical tasks rather than relying only on controlled tests.

Why AI Agents Need Better Evaluation

AI systems are moving beyond answering questions. Increasingly, agents can write code, analyze documents, conduct research and take actions on behalf of users.

However, stronger capabilities also introduce new risks. An agent might take an action without permission, incorrectly report what a user said or claim to have completed a task when it has not.

Traditional benchmarks can measure specific capabilities, but they may not fully capture how AI behaves during complex, real-world interactions. Static tests can also become less useful when models learn to recognize the conditions under which they are being evaluated.

Arena aims to address this challenge by observing complete human-agent workflows, including how tasks develop and what actions an AI system takes along the way.

Its approach uses causal inference methodology to evaluate agent performance across tasks such as coding, document analysis and creative writing. The goal is to produce a more practical view of how AI systems perform when people use them in everyday work.

Arena Launches Its Alignment Index

Alongside the funding announcement, Arena introduced its Alignment Index, a new initiative designed to measure certain AI safety and alignment risks through real-world agent interactions.

The initial preview focuses on three signals.

Unauthorized Action (UA) measures when an AI model takes an action beyond what the user requested or permitted.

False Attribution (FA) identifies cases where an AI system attributes a statement, intention or fact to a user despite evidence provided by that user contradicting the claim.

Deceptive Completion (DC) tracks situations where an AI system tells a user that a task is complete when it has not actually been completed.

Arena says these signals can be checked against actual agent activity and user-provided evidence. The company has released its methodology and initial results for more than 20 frontier AI models.

The Alignment Index is intended to complement Arena’s existing capability rankings. While the leaderboards compare how well AI agents perform tasks, the new index focuses on whether their behavior meets basic expectations around permission, truthful reporting and task completion.

The initial release covers a limited set of measurable signals, and Arena plans to expand the index over time.

Investors Back Arena’s Next Stage of Growth

The Series B was co-led by Lightspeed Venture Partners and Khosla Ventures, with participation from a group of new and existing investors.

The funding gives Arena additional resources to expand its evaluation infrastructure, track more AI models and develop methods for assessing the safety and reliability of increasingly capable AI agents.

As AI adoption spreads across businesses, independent evaluation could become more important for companies deciding which models to use and what tasks to delegate to them.

For developers, evaluation data can help identify weaknesses and areas for improvement. For businesses, it can provide additional evidence when comparing AI tools and assessing potential risks.

Arena’s challenge will be to maintain reliable, transparent evaluation methods as AI systems become more capable and the range of real-world tasks expands.

What Comes Next for Arena?

Arena plans to continue expanding its platform and updating the Alignment Index as it adds new evaluation signals and more model results.

The company is positioning real-world evidence as a central part of AI evaluation, combining large-scale user interactions with methods designed to assess both capability and behavior.

Its $200 million Series B provides funding to pursue that strategy as AI agents take on more complex work.

The long-term test will be whether Arena’s evaluations help developers, businesses and researchers make better decisions about AI performance, safety and reliability.

About Arena

Arena is a San Francisco-based AI evaluation company led by CEO Anastasios Angelopoulos. Its platform evaluates AI models and agents across text, vision, coding, search, video and image tasks.

The company measures performance using real-world interactions and maintains leaderboards covering open and proprietary models. With its new Alignment Index, Arena is expanding its focus from AI capabilities to measurable aspects of safety and alignment.

Arena raised $200 million in Series B funding at a $3.1 billion valuation in a round co-led by Lightspeed Venture Partners and Khosla Ventures.

Recommended Stories for You