Merve Noyan releases slide deck on local AI inference with llama.cpp
AIMerve Noyan has released a slide deck on running AI locally, covering prefill versus decode, MoE versus dense models, VRAM versus unified memory, quantization, and speculative decoding. The deck is built around llama.cpp and is free to reuse with attribution.










