The batching section clicked for me — the prefill vs decode split matters so much more than people realize. Continuous batching is great on the decode side but if you're doing long-context RAG with big retrievals, you're prefill-bound and it's a totally different optimization problem.
KV cache is the other one. PagedAttention got us far but I've been thinking about tiered caching lately — hot KV in HBM, warm in DRAM, cold evicted. TurboQuant-style compression changes the math on what fidelity you actually need at each tier.
Has anyone told you about speculative decoding working well in production? I keep hearing about the latency wins but the wasted compute on rejected tokens seems to scare most teams off before they try it.
Good practical framing. The thing that surprised me most going from single-user to any kind of scale is how completly the bottleneck moves. Solo, I care about single-stream latency, but the moment you batch even 8 requests the whole game becomes memory bandwidth and how full you can keep the batch. A box that felt fast for one user can fall apart at eight. Did the crossover land where you expected for your setup?
The batching section clicked for me — the prefill vs decode split matters so much more than people realize. Continuous batching is great on the decode side but if you're doing long-context RAG with big retrievals, you're prefill-bound and it's a totally different optimization problem.
KV cache is the other one. PagedAttention got us far but I've been thinking about tiered caching lately — hot KV in HBM, warm in DRAM, cold evicted. TurboQuant-style compression changes the math on what fidelity you actually need at each tier.
Has anyone told you about speculative decoding working well in production? I keep hearing about the latency wins but the wasted compute on rejected tokens seems to scare most teams off before they try it.
This made me clearly understand some bottlenecks I'm currently facing and how to overcome them! Always grateful for your support
Enjoyed reading, thanks
Love this!
Good practical framing. The thing that surprised me most going from single-user to any kind of scale is how completly the bottleneck moves. Solo, I care about single-stream latency, but the moment you batch even 8 requests the whole game becomes memory bandwidth and how full you can keep the batch. A box that felt fast for one user can fall apart at eight. Did the crossover land where you expected for your setup?