Agents only improve when you can measure what they do and trust where they run.

This is a technical evening for people building serious AI systems. We will look at two parts of the stack that turn a promising agent into something reliable: evals that show what is working, and sandboxes that give agents safe, isolated environments in which to act.

A YC founder from Abundant will share lessons from building environments and datasets for reinforcement learning, followed by talks and discussion on designing useful evals, comparing agent variants and running experiments safely.

What we will cover

  • Designing evals that measure real agent behaviour

  • Building repeatable evaluation loops for prompts, tools and models

  • Using isolated sandboxes for safe code execution and experimentation

  • Comparing variants without contaminating environments or results

  • Moving from impressive demos to systems that can be tested and improved

Expect focused technical talks, practical examples, audience questions and time to meet other engineers and founders working on AI systems.

Who it is for

Senior software engineers, AI and ML engineers, technical founders, researchers and experienced agent builders. You should be comfortable with technical AI concepts, but you do not need prior experience building eval infrastructure.

Partners

  • Abundant creates environments and datasets for reinforcement learning, drawing on experience in simulation and model training.

  • Geometric AI builds autonomous optimisation for CUDA, Triton and TensorRT, helping ML teams discover and verify faster GPU kernels. We may take photos and short video clips during the evening for event recaps and community updates. If you prefer not to appear, tell one of the hosts when you arrive.

Dublin AI Week

This event is part of Dublin AI Week, a citywide programme of practical AI events by Give(a)Go and Baseline. Explore the full programme.