Foundational articles are always hard because it’s not easy to know when to stop and move forward to the next subject. You struck a nice balance here and left enough crumbs for people to ask their LLM of choice about any gaps.
Hi Miguel , for this Bringing It All Together: The Transformer you have the encode-decoder architecture. is that relevant for understanding transformer. Now the focus is mainly with decoder only architectures I assume.
The Daniel Han's video mentioned another latest finetuning methodology called 'RLVR' (Reinforcement learning by verifiable rewards). Will the articles cover this as well in future?
This Friday, in the hands-on lab we’ll cover distillation, curriculum learning, continued pretraining, etc.! Both the theory and an applied example! :)
Foundational articles are always hard because it’s not easy to know when to stop and move forward to the next subject. You struck a nice balance here and left enough crumbs for people to ask their LLM of choice about any gaps.
This clip lives rent-free in my head:
https://youtu.be/MO0r930Sn_8
Great introduction! Just starting the course, hoping and excited to learn a lot
Great Article.
Hi Miguel , for this Bringing It All Together: The Transformer you have the encode-decoder architecture. is that relevant for understanding transformer. Now the focus is mainly with decoder only architectures I assume.
Hey Miguel,
The Daniel Han's video mentioned another latest finetuning methodology called 'RLVR' (Reinforcement learning by verifiable rewards). Will the articles cover this as well in future?
Yep! The GRPO lesson Will cover that :)
An amazing high level overview!
Thanks for the post. When I started learning about attention this video helped me a lot: https://www.3blue1brown.com/?v=attention
3blue1brown is THE GOAT
Very well organized. Loved the explanation
In 'continued pretraining' - are existing paremeters frozen and new layers added and trained on?Is the concept similar to distilled models?
This Friday, in the hands-on lab we’ll cover distillation, curriculum learning, continued pretraining, etc.! Both the theory and an applied example! :)