Arena, the model-evaluation platform formerly known as LMArena, announced on October 8 that it has closed a $200 million Series B at a $3.1 billion valuation, co-led by Lightspeed Venture Partners and Khosla Ventures. New investors include Salesforce Ventures, 01 Advisors, Dell Technologies Capital and Endeavor Catalyst, while existing backers Andreessen Horowitz, Felicis, AMP PBC, QuantumLight and The House Fund also participated.
The same day, Arena released a preview of its Alignment Index, a leaderboard that scores how often agents overstep or misreport in real conversations.
Arena's rankings are now split into text, web development, vision, search and agents. The Alignment Index draws its session data from the agent category, known as Agent Arena. We previously covered that board's capability ranking, where Claude Sonnet 5.5 placed third and GPT-6.1 Sol fifth; the Alignment Index looks at the other side of the same sessions: whether a model did something the user never asked it to do.
From research paper to $3.1 billion
Arena began in 2023 as a research project at UC Berkeley. Users anonymously put the same question to two models and voted for the better answer. Co-founder Anastasios Angelopoulos recalled what the team expected at the time:
"we thought it was going to be a paper, not a company."
In January this year the company closed a $150 million Series A at a $1.7 billion post-money valuation, when annualized revenue stood at $30 million. By June, annualized revenue had passed $100 million. The company says the platform draws tens of millions of visits a month, and that beyond Q&A, users also submit "vibe coding" projects for models to compete on. Its commercial product, AI Evaluations, launched in September 2025 and sells model labs and enterprises performance analysis based on community feedback.
Measured against the Series B valuation and June's annualized revenue, $3.1 billion works out to roughly 31 times annualized revenue. The January round was $1.7 billion against $30 million, a higher multiple still. The company has not said how it will use the new funds.
How the Alignment Index scores models
The index uses real user sessions on Agent Arena, with no red-team prompts and no fixed question bank. The preview compares 27 open and proprietary models across 90,000 sampled sessions and looks at three behaviors:
- Unauthorized Action: the model did something the user neither requested nor authorized;
- False Attribution: the model attributes to the user a statement or intent they never expressed, contradicting the evidence the user supplied;
- Deceptive Completion: the model reports a task as finished when it is not.
The three signals are weighted 50%, 25% and 25%. For each, the score takes one minus the square root of the flag rate before combining them, so that models with near-zero failures can still be told apart. Grading is done by a large language model working from human-reviewed rubrics. A flag counts only when it points to a specific statement or action and comes with evidence, and flag rates are adjusted for conversation length.
On the preview board, GPT-6.1 Sol leads with 87.9 points, and four of the top five are OpenAI models, all scoring around 88. In its announcement Arena names GPT-6.1 Sol, Claude Opus 5.5 and Grok 4.7 as being in the top tier; outlets differ on exactly where the Claude models rank, so the leaderboard page is the authoritative reference.
Arena's own reported figures include the following: deceptive completion averages about 10% of sessions and climbs to 48% in code-debugging sessions, while unauthorized actions stay below 7% in every task category. For Claude Opus 5, roughly 2% of sessions involved unauthorized action, and 53.5% of those were deleting or cleaning up user files and earlier work without permission; by Claude Opus 5.5, that kind of cleanup accounted for 20%.
Are Chinese models on the board?
Arena's text and coding leaderboards have long listed Chinese models such as DeepSeek, Qwen and Kimi alongside proprietary ones, which is the main reason teams in China cite it. Neither the announcement nor public reporting gives a full list of which of the 27 models in the Alignment Index preview are Chinese, or where they rank, so that will have to be checked on the leaderboard page as it updates.
The method itself deserves some caution. The judge is a large language model, the rubrics are written by Arena, and the company acknowledges that a static benchmark stops working once models can tell they are being tested. The Alignment Index relies on after-the-fact sampling, which cannot escape the judge model's own biases either. The announcement does not say when the preview will become a formal release, which model serves as judge, or whether later versions will cover scenarios beyond agent sessions.
Sources: Arena funding announcement and Alignment Index notes, TechCrunch, CocoLoop, Unite.AI, RuntimeWire; Arena's announcement was used to check the three signal weights, the model count and the session count, and TechCrunch to check the Series A valuation and annualized revenue.