0:00
/

Paid episode

The full episode is only available to paid subscribers of The Neural Maze

Finetuning Sessions – Week 5 Office Hours

RL for LLM Alignment with PPO, DPO & KTO

Here's the recording from yesterday's live session, covering the Week 5 content of the Finetuning Sessions.

During office hours, we focused on:

  • Why SFT hits a ceiling and how RLHF bridges the gap between fluent output and human-preferred output

  • The three-step RLHF process: preference collection, reward modeling, and policy optimization

  • How reinforcement lea…

User's avatar

Continue reading this post for free, courtesy of Miguel Otero Pedrido.