Two UC Berkeley researchers surveyed teams deploying agents across dozens of application domains, then paired the results with in-depth case studies of how those systems are actually built, evaluated, and operated.
The findings are unusually concrete.
Melissa Pan and Negar Arabzadeh, co-lead authors of Measuring Agents in Production, will present the work jointly in a twenty-minute talk.
Then we test the implication.
If reliability is the bottleneck, self-improvement looks like the obvious next move: let an agent optimize itself against an evaluation. But evaluations are proxies, and proxies can separate from the outcomes they are meant to represent. In retrieval systems, the configuration that produces the best ranked list is often not the one that produces the best final answer.
Point an automated improvement loop at the wrong proxy and the failure does not merely persist. It compounds.
Vignesh Baskaran builds self-improving agent systems in the open. He will put that problem to both researchers onstage—and put the assumptions behind his own work under the same scrutiny. Then the room joins the argument.
Short presentation. Long discussion. Paper circulated in advance.
The Frontier Research Club is a curated forum for rigorous technical discussion at the frontier of AI. We bring together researchers from frontier AI labs, Stanford, Berkeley, and the teams deploying these systems in production to examine concrete work: papers, methods, experiments, and results.
The emphasis is on assumptions, evaluation design, failure modes, and what counts as convincing evidence. Presentations stay brief so that questions, critique, and discussion can take up most of the room.
Agenda
5:30–6:30 PM — Doors, dinner, and conversation6:30–8:00 PM — Research presentation, fireside, and room discussion8:00–8:30 PM — Closing conversation
The Research
Measuring Agents in Production
Presented jointly by Melissa Pan and Negar Arabzadeh
For all the attention around agentic AI, remarkably little has been published about how production systems are actually built. Most of that knowledge remains inside the companies deploying them.
Measuring Agents in Production offers the first systematic view: a large-scale survey paired with in-depth case studies across dozens of application domains. The work examines how much autonomy deployed agents actually have, when teams fine-tune and when they do not, how these systems are evaluated, and what fails first once real users depend on the output.
Melissa Pan is a PhD student in computer science at UC Berkeley’s Sky Computing Lab and lead author of Measuring Agents in Production. The paper was selected for an oral presentation at ICML 2026—one of 168 papers selected from 24,661 submissions.
She also co-authored MAST, a widely used taxonomy of multi-agent failure modes, and Fantastic Adaptive Taxonomies and How to Use Them, which received Best Paper at the FAGEN workshop at ICML 2026. Melissa previously worked at Google and IBM and holds an MS in Electrical and Computer Engineering from Carnegie Mellon University.
Negar Arabzadeh is a postdoctoral researcher at UC Berkeley EECS working with Professor Matei Zaharia and a co-lead author of Measuring Agents in Production. She received her PhD in Computer Science from the University of Waterloo and was awarded the university’s 2026 Faculty of Mathematics Doctoral Prize.
Her research spans information retrieval, retrieval-augmented generation, and agentic search. She co-authored DeepScholar-base, an open-source deep-research pipeline competitive with proprietary systems while operating at roughly twice the speed. Her previous research experience includes Microsoft Research, Spotify Research, and Google. She is also Head of Data Science at Reviewerly.
The Fireside
What Changes When the Agent Can Improve Itself?
Vignesh Baskaran in conversation with Melissa Pan and Negar Arabzadeh
Vignesh Baskaran is co-founder and CTO of Hexo Labs, where he built SIA, an open-source framework that updates both the agent harness and the model weights of a task-specific agent within a single self-improvement loop.
He will press the researchers on what their production findings imply for automated improvement: what should be optimized, which evaluations can be trusted, and what happens when the metric quietly diverges from the outcome.
The researchers will have the opportunity to press back.
Want to present your work?
If you have a research paper you’d like to discuss at one of our next sessions, please submit it for consideration. Submit your paper here!
Who should attend
-
Researchers working on agent evaluation, retrieval, reliability, or self-improvement
-
Engineers operating agent systems where real users depend on the output
-
Teams building evaluation infrastructure, observability, and agent tooling
-
Anyone who has shipped an agent and watched it fail in a way no benchmark predicted
-
Anyone who suspects their metrics and their actual goals have quietly come apart
Capacity is limited. Registration is subject to approval.
We will take photos and short video clips for event recap and promotion. By attending, you consent to being photographed and recorded, and to the use of those images and clips by the organizers on social media and other event marketing channels.
🌐 Connect with Frontier Research Club
-
Luma Calendar: luma.com/frontiersyndicate
-
YouTube: youtube.com/@FrontierResearchClub
-
Instagram: @frontierresearchclub
-
Email: [email protected]
Hosted by
The Frontier Syndicate is a research-native venture network connecting frontier technology researchers, builders, and investors through curated convenings and early-stage capital. Across the Bay Area, we host recurring research forums, builder nights, and private investor dinners—and back exceptional companies emerging from the technical networks we convene.
Hexo Labs is a research lab building recursive self-improving AI. Its open-source SIA framework updates both the harness and model weights of a task-specific agent within the same self-improvement loop, producing state-of-the-art results across multiple domain benchmarks. Hexo also supports the broader research community through grants and direct collaboration on difficult problems in science and engineering.