Here's the recording from yesterday's live session, covering the Week 5 content of the Finetuning Sessions.
During office hours, we focused on:
Why SFT hits a ceiling and how RLHF bridges the gap between fluent output and human-preferred output
The three-step RLHF process: preference collection, reward modeling, and policy optimization
How reinforcement lea…














