Writing

Optimizing Local LLMs on Apple Silicon / EP02

Why Even a 128GB Mac Can Run Out of Memory

Model weights are only part of a 128GB Mac’s memory budget. I mapped out KV caches, runtime allocations, development tools, and which models might stay loaded.

In this post

Surely 128GB buys me some breathing room

The number that caught my eye when ordering the Mac Studio was 128GB. My current iMac has 8GB. Surely that would mean fewer interruptions over memory, at least for a while.

In EP01, I split slow LLM responses into prefill and decode. This time the question comes even earlier: will everything fit at all?

At first I looked at model download sizes. But chatting with one big model isn’t all I want to do. I’d like a language model available for Hermes, plus image and audio models when needed. I’ll also be building apps and running simulators beside them.

Written out like that, I’m asking quite a lot of those 128GB. The Mac Studio is still on order. This is a note about what to check when it arrives, rather than a report that I’ve already run out of memory.

The model file needs company

A model file mainly holds weights and data required by its storage format. Running it also needs a KV cache, intermediate results, temporary working space and runtime allocations. The Transformer inference analysis distinguishes weights from data used during execution.

Separating the budget makes it easier to see where to leave room. These numbers illustrate one allocation of 128GB; they are not requirements calculated for selected models.

One way to budget 128GB

Illustrative budget · not measured scenario comparisons

Illustrative 128GB unified-memory budgetmacOS and dev tools: 20GB (15.6%); Base LLM weights, etc.: 28GB (21.9%); Vision, image and audio models: 32GB (25.0%); KV cache and working space: 16GB (12.5%); Headroom: 32GB (25.0%)2028321632
The same budget in numbers
CategoryGB%
macOS and dev tools2015.6
Base LLM weights, etc.2821.9
Vision, image and audio models3225.0
KV cache and working space1612.5
Headroom3225.0
Total128100

This graph is an illustrative memory budget, not a measurement. Actual usage depends on the model, context length, runtime and concurrent work.

The 32GB headroom is capacity intended to remain unused, not reserved or occupied memory. Categories are not fixed requirements. OS, weights, KV cache and runtime cache can interact; measured counters must not simply be added together.

How the task changes the plan

Expand each scenario for its setup and allocation principle. The numbers above are not A/B/C requirements. Models, quantization, context and runtimes are undecided; numerical scenario comparisons await measurements after delivery.

A · Everyday operation

Hermes · medium general LLM · JevMLX · dev tools · small speech model

Keep frequent language and decision models resident first, with room for development. Check Hermes and tool overhead too.

B · Multimodal work

Medium general LLM · vision · image generation · audio models · dev tools

Load models in task order rather than all at once. Include peak activation and working-buffer memory during generation.

C · Large LLM run

Large Qwen or DeepSeek candidate · weights and KV cache · macOS/runtime · minimal other resident models

Reduce other models first, then check total weights, context and working space. MoE active parameters alone do not establish whether it fits.

Apple Silicon’s unified memory is shared by CPU and GPU. MLX explains that both can work on the same arrays. Avoiding transfers into separate VRAM helps, but macOS and other apps need space in that same memory too.

“128GB Mac, therefore a 128GB model” leaves things out. Weights aren’t the whole budget, and the system needs headroom. Metal also exposes a recommended working-set size. I won’t treat installed capacity as the amount safely available to GPU work.

I also need to distinguish dense models from MoE. A dense model generally uses most of its weights for each token; MoE selects some experts. MoE reports therefore separate total parameters from active parameters involved in each token’s computation. Fewer active parameters don’t remove the memory cost of all expert weights if those weights remain resident.

Streaming the required experts from SSD can reduce residency, but adds I/O to read and move them. Whether a usable implementation exists for my Mac is a separate question. I’ll explore this in EP04.

Longer conversations need room beyond the weights

A KV cache keeps previously calculated keys and values for reuse. For conventional full attention, more retained tokens mean more stored data. The cache implementation may grow allocations as needed or reserve a maximum size ahead of time.

For a simple case, the payload is approximately:

KV bytes ≈ 2 (K and V) × layers × KV heads × head dimension × retained tokens × bytes per element

The diagram assumes 32 layers, 8 KV heads, head dimension 128, two bytes per element and one session. This is arithmetic for an imaginary configuration, not a particular Qwen model or a measurement from my ordered Mac. Buffer headroom and metadata are excluded.

Calculated full-attention KV payload: 4096 tokens need 512MiB, 8192 need 1GiB and 16384 need 2GiB in an assumed configuration
Weights can stay unchanged while conversation state grows. Tokens here include retained input and generated output. These are theoretical KV payload sizes, not total model memory.

This formula isn’t universal. GQA shares KV heads across query heads; MQA shares a single KV head. With other conditions equal, fewer KV heads mean less stored data. Sliding-window layers can limit cache growth, and hybrid models can use different state structures across layers. I need the actual configuration, not just the model’s name.

Four-bit weights don’t automatically mean a four-bit KV cache either. Cache quantization has its own support and settings. FP16 uses two bytes per element and FP8 one, but storage may also need data such as scales. The vLLM FP8 KV guide documents supported combinations; that does not establish support in a Mac runtime. Download size alone won’t tell me whether long conversations will fit.

Two models and two conversations are different problems

Two conversations with one model don’t necessarily require two copies of its weights. An engine such as the llama.cpp server can serve multiple requests. Per-session state and cache space still matter. Shared prefixes and allocation policies depend on the implementation.

Launching separate model processes doesn’t automatically provide the sharing of one engine. Two different models also bring their own weights and working allocations.

What increases Memory to check
One conversation’s length KV cache and other conversation state
Concurrent sessions Session state, batch workspace and actual sharing
Different loaded models Each model’s weights, caches and working space
Dev tools and simulators System headroom and memory pressure

I’m also curious about long Hermes conversations containing tool results. A short question from me might become a much longer model input. I’ll need to record the tokens actually sent.

Do all the models need to stay loaded?

I’m considering a medium general LLM and a small decision model such as JevMLX for frequent use, with STT or embeddings added as needed. Large LLM candidates such as Qwen3.8 Flash Next, plus image, video and music generation models, would load on demand. Candidate names, actual versions and memory requirements still need confirmation. I don’t plan to keep them all loaded at once.

Image and video generation need more than room for weights. Activations and working buffers can grow during inference, and task conditions such as resolution or frame count affect the peak. I’ll use the Diffusers memory guide as a reference, then check the runtime I actually use. Large tasks may require unloading another model first.

Planned workflow: keep a frequent language model resident, check headroom for an image or audio request, load the model and release it after work
A workflow I'd like to build, not memory management Hermes already performs. Waiting or unloading a resident model are options to evaluate against actual delays.

Saving memory means the first request may wait for loading. Whether I unload immediately or keep a model around briefly will depend on frequency and load time. A file on SSD isn’t the same as a model ready in memory.

This is still an operating plan. Alongside “how many can stay loaded?”, I want to know “how long does bringing one back take?”

Read the labels before adding memory numbers

I plan to use Activity Monitor alongside runtime counters. Apple’s guide says memory pressure reflects free memory, swap rate, wired memory and file cache. Rather than only checking how much memory is used, I want to track memory pressure, swap and any response slowdown as conversations get longer.

MLX exposes active memory, peak memory and allocator cache. Allocator cache holds space for later allocations; it isn’t the conversation’s KV cache.

Adding process memory to MLX counters can double-count space. MLX’s peak isn’t the whole system’s unified-memory peak either. OS file cache left after unloading also needs a separate interpretation.

When it arrives, change one thing at a time

Launching several models and a simulator immediately would make changes hard to explain. I’ll fix model, quantization and runtime versions, then work through these conditions.

Order Change Record alongside it
1 Input length with one model Actual tokens, KV settings, memory and response changes
2 Concurrent sessions for that model Per-session length, batch settings, delays and sharing
3 Add a second model Load time, actual concurrency, peak and post-release state
4 Add dev tools and simulators Quiet conditions versus my normal development session

I’d like to capture loading, first response, generation, idle time and release separately. Alongside EP01’s TTFT, prefill and decode metrics, I’ll watch pressure, compression and swap changes. Keeping app development comfortable matters more to me than deliberately pushing the machine to its limit.

What I haven’t measured

I don’t yet know which models will fit comfortably, or what happens with image and audio models alongside them. The arithmetic is arithmetic; the diagrams describe a plan. Actual usage and speed have to wait for the machine.

Initially I wanted to load the biggest model possible. Now I’d be happy to leave a useful everyday model running while building apps comfortably, and bring out the large one when needed.

Even if everything fits, slow answers will raise another question. In EP03 I’ll look at how speculative decoding and MTP try to shorten that wait.