if you want to get every last bit of performance out of your local setup, you need to know a bit about inference 🥵
but we've got you covered, shipping conceptual guides 🔥
here's the first one about prefill, decode, KV cache what you should optimize for 🙌🏻
