Daytona AI Researchers - Stanford, October 6, 2026
On Tuesday, October 13, Daytona and FounderCoHo are again co-hosting an exclusive, high-signal evening dedicated to researchers at Stanford University to explore when we take long-horizon, stateful agents seriously — from the infrastructure that makes them possible, to the evaluation frameworks that make them trustworthy.
Agenda
🕒 5:30 pm – 6:00 pmWelcome Reception and Opening Remarks
🎤 Marijan Cipcic, Principal Events Manager at Daytona
🕒 6:00 pm – 6:15 pmTalk "Today's Agents Don't Live In Episodes"
🎤 Muhammad Annas Hashmi, DevRel at Daytona
Outline:
The 'episode' (short, stateless, resettable) has been RL's foundational abstraction since ATARI. It underpins the Gym API, GRPO, PPO, and the conventional sandbox lifecycle. Today's agents no longer fit it. Tasks span for days; the env state at hour 18 of an agent session with warm caches, installed deps, live processes, open sockets, dirty git tree, is worth hours of wall clock to reproduce.Three things are scaling simultaneously. Rollout horizon: seconds -> days. Env state: disposable between episodes -> first-class learning substrate. Branching: absent in modern LLM-RL -> speculative fork trees. Each stresses the inherited toolkit in a different way, and all three have been gated on the same missing primitives: VMs you can fork cheaply, pause without killing processes, snapshot mid-run, and resume hours later.This talk walks through what opens up when those primitives become available. Live demo of long-horizon sessionful rollouts, mid-trajectory forking, and cross-calendar-time training. The research questions that follow (long-horizon benchmarks, speculative RL algorithms, event-driven training, to name a few) are where the next wave of agent RL gets built.
🕒 6:15 pm – 6:30 pmTalk "Scaling Automatic Research Agents via World Models"
🎤 Dr. Zhenyu Liao, Senior applied scientist in Amazon working on Reinforcement learning on VLM and AutoResearch Agent, Efficient World Action Model
Outline:
AutoResearch agents can automate empirical research, but scaling RL is bottlenecked by costly environment execution. So we proposed World Model RL (WMRL) replaces real execution with a learned world model to accelerate training. To address imperfect world-model rewards, WMRL uses Online Debiasing to correct bias and Inverse-Variance Denoising to suppress noise, with theoretical improvements in convergence. Empirically, WMRL achieves 3–4× faster training while outperforming standard RL baselines, and its 4B/9B agents outperform much larger 48B/120B agents on held-out benchmarks. The method also transfers successfully to post-training embodied VLA policies, suggesting that WMRL can generalize beyond AutoResearch.
🕒 6:30 pm – 6:45 pmTalk "RL Environments: The Next Frontier of LLM Training"
🎤 Dr. Xide Xia, Co-Founder at Steadyworks (AI data training and RL research lab); previously Staff Research Scientist at Meta, core contributor to Llama 3
Outline:
As LLMs move from generating answers to executing complex, multi-step tasks, RL environments are becoming increasingly important for training. Human-labeled data alone is not enough: models need to learn through interaction, feedback, and trial and error. This talk will discuss why RL environments matter and what makes them effective for training next-generation Al agents.
🕒 6:45 pm – 7:00 pmTalk "The Next Hill: What Evals Need to Measure Next"
🎤 Xiangyi Li, Founder at BenchFlow. Author of SkillsBench
Outline:
Benchmarks are saturating fast, especially in domains with stable fundamentals (finance, law) or easily verified results (math, coding). What stays hard is human interaction, company context, and fields like AI and science where new knowledge keeps appearing. This talk traces how eval formats evolved, from MMLU to SWE-bench and PaperBench. It then argues that the next wave of evals will cover science and AI research, and will measure alignment and safety, agent behavior (over-eager, under-eager, reckless), and multi-agent and physical-world tasks. Examples come from BenchFlow's BenchGuard and FrontierPhysics.
🕒 7:00 pm - 8:30 pm
Networking
With food and beverages
About event
An engaging meetup designed for AI researchers to connect, share ideas, and explore the latest advancements in artificial intelligence. The event features informal networking, short talks, and discussions on current research trends, fostering collaboration and knowledge exchange within the AI community.