0:00
/

Paid episode

The full episode is only available to paid subscribers of The Neural Maze

Finetuning Sessions – Week 6 Office Hours

Reward Functions, Reward Hacking & GRPO Variants (DAPO, GSPO, Dr. GRPO)

Here's the recording from yesterday's live session, covering the Week 6 content of the Finetuning Sessions.

During office hours, we focused on:

  • Why PPO's four-model setup becomes an infrastructure bottleneck and how GRPO eliminates the critic

  • How GRPO works: group-based advantage estimation and grading on a curve

  • The difference between GRPO and RLVR — optim…

User's avatar

Continue reading this post for free, courtesy of Miguel Otero Pedrido.