0:00
/

Paid episode

The full episode is only available to paid subscribers of The Neural Maze

Finetuning Sessions – Week 4 Office Hours

Covering QLoRA, GPU memory and quantization techniques

Here's the recording from yesterday's live session, covering the Week 4 content of the Finetuning Sessions.

During office hours, we focused on:

  • The VRAM bottleneck in LLM training and why memory — not compute — is the real constraint

  • How number formats evolved from FP32 to 4-bit, and why NF4 is purpose-built for neural network weights

  • The difference between

User's avatar

Continue reading this post for free, courtesy of Miguel Otero Pedrido.