Most teams still evaluate their AI the same way: run a few prompts, eyeball the outputs, and ship if it feels right.
That works until it doesn't, until the agent fails silently in production, the hallucination reaches a customer, or nobody can explain why last week's change made things worse.
This fireside gets into what comes after the vibe check: what a production eval actually is and how to build one, where eval datasets come from, using online evals as live guardrails, the LLM-as-judge vs. purpose-built eval models debate, and whether you can grade a multi-step agent's whole trajectory or just its final answer. Plus the practical calls every team faces, build vs. buy, and who owns evals inside a company.
🎤 🎤 🎤 SPEAKERS 🎤 🎤 🎤
-
LIAM BUSH | Deployed Engineer @ LangChain LangChain is the default open-source framework for building with LLMs, 118,000+ GitHub stars, $1.25B valuation, powering agents and eval pipelines for thousands of companies, and the team behind LangSmith for evals and observability.
-
LOTTE VERHEYDEN | Head of Developer Relations @ Langfuse Langfuse is the leading open-source platform for LLM observability and evals, 26M+ SDK installs a month, trusted by 19 of the Fortune 50, and acquired by ClickHouse in early 2026. The open-source, self-hostable tool that became a developer favorite for measuring LLM quality.
-
SOUMYA MOHAN | Head of Product @ GalileoGalileo is an evaluation platform built to make AI agents reliable, catching hallucinations and failures before production. $68M raised, 834% revenue growth in a year, with research-backed metrics that score AI output automatically, at scale.
-
EMMANUEL TURLAY | Director of Engineering @ CoreWeave CoreWeave is the AI Cloud powering the world's top AI labs, with 2025 revenue guided near $5B. Its Weights & Biases platform brings rigor to LLM evaluation, with Weave monitoring production AI agents in real time across any cloud.
-
BRADEN HOLSTEGE | VP, Enterprise AI @ MercorMercor works with frontier AI labs and leading enterprises to evaluate and improve AI systems. Mercor's Enterprise team applies that expertise inside companies, building evaluations around real-world workflows and customer-specific standards to help move AI from prototype to production.
🏠 HOSTED BY FORWARD DEPLOYED
Forward Deployed is an AI-Native startup building agentic products for PE portfolio companies, Insurers, and other Enterprise teams. We also host fireside chats and podcasts with operators at the forefront of forward deployed engineering.
📅 Tue, Sep 1st, 6-8:30pm PT📍 Super secret, San Francisco 🎟 Seats are limited
✌️ Hosted at the Founders Cafe, AngelList's founder community and co-working space on the first floor of AngelList HQ. Apply here.