apple-silicon
- Why Even a 128GB Mac Can Run Out of Memory
Model weights are only part of a 128GB Mac’s memory budget. I mapped out KV caches, runtime allocations, development tools, and which models might stay loaded.
- Why Are Local LLMs Slow? — Prefill and Decode
Why can an LLM take ages to start, then stream quickly? Prefill, decode, TTFT, KV caches, and a benchmark plan for the Mac Studio I am waiting for.