All topics

llm-inference

  1. Why Even a 128GB Mac Can Run Out of Memory

    Model weights are only part of a 128GB Mac’s memory budget. I mapped out KV caches, runtime allocations, development tools, and which models might stay loaded.

  2. Why Are Local LLMs Slow? — Prefill and Decode

    Why can an LLM take ages to start, then stream quickly? Prefill, decode, TTFT, KV caches, and a benchmark plan for the Mac Studio I am waiting for.