12 Comments
User's avatar
Marcelo Acosta Cavalero's avatar

Foundational articles are always hard because it’s not easy to know when to stop and move forward to the next subject. You struck a nice balance here and left enough crumbs for people to ask their LLM of choice about any gaps.

This clip lives rent-free in my head:

https://youtu.be/MO0r930Sn_8

Victor Adewoyin's avatar

Great introduction! Just starting the course, hoping and excited to learn a lot

Ankana Mukherjee's avatar

Great Article.

Manjunath's avatar

Hi Miguel , for this Bringing It All Together: The Transformer you have the encode-decoder architecture. is that relevant for understanding transformer. Now the focus is mainly with decoder only architectures I assume.

manishlearnsai's avatar

Hey Miguel,

The Daniel Han's video mentioned another latest finetuning methodology called 'RLVR' (Reinforcement learning by verifiable rewards). Will the articles cover this as well in future?

Miguel Otero Pedrido's avatar

Yep! The GRPO lesson Will cover that :)

manishlearnsai's avatar

An amazing high level overview!

jose's avatar

Thanks for the post. When I started learning about attention this video helped me a lot: https://www.3blue1brown.com/?v=attention

Miguel Otero Pedrido's avatar

3blue1brown is THE GOAT

Santhanalakshmi's avatar

Very well organized. Loved the explanation

tanzeel's avatar

In 'continued pretraining' - are existing paremeters frozen and new layers added and trained on?Is the concept similar to distilled models?

Miguel Otero Pedrido's avatar

This Friday, in the hands-on lab we’ll cover distillation, curriculum learning, continued pretraining, etc.! Both the theory and an applied example! :)