Most teams still evaluate their AI the same way: run a few prompts, eyeball the outputs, and ship if it feels right.

That works until it doesn't, until the agent fails silently in production, the hallucination reaches a customer, or nobody can explain why last week's change made things worse.

This fireside gets into what comes after the vibe check: what a production eval actually is and how to build one, where eval datasets come from, using online evals as live guardrails, the LLM-as-judge vs. purpose-built eval models debate, and whether you can grade a multi-step agent's whole trajectory or just its final answer. Plus the practical calls every team faces, build vs. buy, and who owns evals inside a company.

🎤 🎤 🎤 SPEAKERS 🎤 🎤 🎤

  • LIAM BUSH | Deployed Engineer @ LangChain LangChain is the default open-source framework for building with LLMs, 118,000+ GitHub stars, $1.25B valuation, powering agents and eval pipelines for thousands of companies, and the team behind LangSmith for evals and observability.

  • LOTTE VERHEYDEN | Head of Developer Relations @ Langfuse Langfuse is the leading open-source platform for LLM observability and evals, 26M+ SDK installs a month, trusted by 19 of the Fortune 50, and acquired by ClickHouse in early 2026. The open-source, self-hostable tool that became a developer favorite for measuring LLM quality.

  • SOUMYA MOHAN | Head of Product @ GalileoGalileo is an evaluation platform built to make AI agents reliable, catching hallucinations and failures before production. $68M raised, 834% revenue growth in a year, with research-backed metrics that score AI output automatically, at scale.

  • EMMANUEL TURLAY | Director of Engineering @ CoreWeave CoreWeave is the AI Cloud powering the world's top AI labs, with 2025 revenue guided near $5B. Its Weights & Biases platform brings rigor to LLM evaluation, with Weave monitoring production AI agents in real time across any cloud.

  • BRADEN HOLSTEGE | VP, Enterprise AI @ MercorMercor works with frontier AI labs and leading enterprises to evaluate and improve AI systems. Mercor's Enterprise team applies that expertise inside companies, building evaluations around real-world workflows and customer-specific standards to help move AI from prototype to production.

🏠 HOSTED BY FORWARD DEPLOYED

Forward Deployed is an AI-Native startup building agentic products for PE portfolio companies, Insurers, and other Enterprise teams. We also host fireside chats and podcasts with operators at the forefront of forward deployed engineering.

📅 Tue, Sep 1st, 6-8:30pm PT📍 Super secret, San Francisco 🎟 Seats are limited

✌️ Hosted at the Founders Cafe, AngelList's founder community and co-working space on the first floor of AngelList HQ. Apply here.