Have you ever wondered what goes into making an LLM? Wonder no more, as we will delve into the birds and the bees on how a baby LLM is born and grows into something that can eat the world!
Agenda: 17:00 Arrive, hang out, get some snacks, beer, soda, 18:15 Intro 18:30 The Hitchhiker’s Guide to European Portuguese LLMs Duarte Carmo - Independent AI/ML research engineer and contractor
Large language models are eating the world. Frontier labs keep pushing the boundary, open-weights models are quickly closing the gap, Europe is — as always — stuck somewhere in the middle, and even Portugal has now released AMÁLIA, its own effort in the space.
But what does it actually take to build an LLM trained on a niche dataset? Why would you? And where do you start when there is so little data? In this talk, I’ll walk through my work on European Portuguese LLMs: from building benchmarks, to finding and filtering data, to training models, and everything in between. Even though this talk focuses on Portuguese, it should be easily applicable to any language/domain!
19:30 More socializing, discussions, hanging out 22:00 Doors close
Sponsors: Thanks to NumFocus for running Pydata!
Special thanks to PROSA for hosting us!