ollama pull qwen2.5:72b then Activity Monitor lied for 40 min
Tried the "Mac Studio runs 70B easy" blog again on an M2 Ultra 64GB. ollama pull qwen2.5:72b-instruct-q3_K_M. Activity Monitor peaked ~48GB, then settled at 22GB while tokens crawled at ~3/s with num_ctx 8192. Disk was thrashing the whole time.
Switched back to qwen2.5:14b-instruct-q4_K_M, same prompt, num_ctx 16384 — ~28 tok/s and Cursor didn't freeze when I alt-tabbed. The 72b run also left two keep_alive ghosts until I ollama stop twice.
Unpopular but: if your demo needs half the weights in swap, just pay the $20/mo cloud seat. 14b + tight context still beats a half-loaded giant on this machine.
2 comments
Join the discussion
Log in to comment.
same here, sorry for english. default num_ctx on ollama was killing my demos for weeks — model looked "loaded" but kv cache was tiny and answers got weird after ~2k tokens.
i pin
num_ctxin Modelfile now. 14b q4 with 16k is more useful than 70b that swaps. $20 cloud for long context is still cheaper than another Mac for me.I want wall-clock numbers before I trust either side. 3 tok/s with swap vs 28 tok/s on 14b is the right comparison — but did the 72b run ever finish a full agent turn, or did you kill it mid-tool-call?
Also: keep_alive ghosts are a process-management bug, not a model-size argument. If ollama leaves orphans after stop, that belongs in the issue tracker, not in the "buy cloud" conclusion.