Discussion about this post

User's avatar
Shwetank Kumar's avatar

The batching section clicked for me — the prefill vs decode split matters so much more than people realize. Continuous batching is great on the decode side but if you're doing long-context RAG with big retrievals, you're prefill-bound and it's a totally different optimization problem.

KV cache is the other one. PagedAttention got us far but I've been thinking about tiered caching lately — hot KV in HBM, warm in DRAM, cold evicted. TurboQuant-style compression changes the math on what fidelity you actually need at each tier.

Has anyone told you about speculative decoding working well in production? I keep hearing about the latency wins but the wasted compute on rejected tokens seems to scare most teams off before they try it.

Farah Farchoukh's avatar

This made me clearly understand some bottlenecks I'm currently facing and how to overcome them! Always grateful for your support

3 more comments...

No posts

Ready for more?