Have you ever wondered what goes into making an LLM? Wonder no more, as we will delve into the birds and the bees on how a baby LLM is born and grows into something that can eat the world!

Agenda: 17:00 Arrive, hang out, get some snacks, beer, soda, 18:15 Intro 18:30 The Hitchhiker’s Guide to European Portuguese LLMs Duarte Carmo - Independent AI/ML research engineer and contractor

Large language models are eating the world. Frontier labs keep pushing the boundary, open-weights models are quickly closing the gap, Europe is — as always — stuck somewhere in the middle, and even Portugal has now released AMÁLIA, its own effort in the space.

But what does it actually take to build an LLM trained on a niche dataset? Why would you? And where do you start when there is so little data? In this talk, I’ll walk through my work on European Portuguese LLMs: from building benchmarks, to finding and filtering data, to training models, and everything in between. Even though this talk focuses on Portuguese, it should be easily applicable to any language/domain!

19:30 More socializing, discussions, hanging out 22:00 Doors close

Sponsors: Thanks to NumFocus for running Pydata!

Special thanks to PROSA for hosting us!

  • ai