a16z general partner Anjney Midha sits down with LMArena cofounders Anastasios N. Angelopoulos, Wei-Lin Chiang, and Ion Stoica to talk about the future of AI evaluation. As benchmarks struggle to keep up with the pace of real-world deployment, LMArena is reframing the problem: what if the best way to test AI models is to put them in front of millions of users and let them vote? The team discusses how Arena evolved from a research side project into a key part of the AI stack, why fresh and subjective data is crucial for reliability, and what it means to build a CI/CD pipeline for large models. They also explore: - Why expert-only benchmarks are no longer enough - How user preferences reveal model capabilities — and their limits - What it takes to build personalized leaderboards and evaluation SDKs - And why real-time testing is foundational for mission-critical AI Chapters: 00:00:04 - LLM evaluation: From consumer chatbots to mission-critical systems 00:06:04 - Style and substance: Crowdsourcing expertise 00:18:51 - Building immunity to overfitting and gaming the system 00:29:49 - The roots of LMArena 00:41:29 - Proving the value of academic AI research 00:48:28 - Scaling LMArena and starting a company 00:59:59 - Benchmarks, evaluations, and the value of ranking LLMs 01:12:13 - The challenges of measuring AI reliability 01:17:57 - Expanding beyond binary rankings as models evolve 01:28:07 - A leaderboard for each prompt 01:31:28 - The LMArena roadmap 01:34:29 - The importance of open source and openness 01:43:10 - Adapting to agents (and other AI evolutions)