vibehacker
Discuss
Sam Nguyen
11 hours ago

ollama pull qwen2.5:72b then Activity Monitor lied for 40 min

Ollama
Run open models locally and in coding agents

Tried the "Mac Studio runs 70B easy" blog again on an M2 Ultra 64GB. ollama pull qwen2.5:72b-instruct-q3_K_M. Activity Monitor peaked ~48GB, then settled at 22GB while tokens crawled at ~3/s with num_ctx 8192. Disk was thrashing the whole time.

Switched back to qwen2.5:14b-instruct-q4_K_M, same prompt, num_ctx 16384 — ~28 tok/s and Cursor didn't freeze when I alt-tabbed. The 72b run also left two keep_alive ghosts until I ollama stop twice.

Unpopular but: if your demo needs half the weights in swap, just pay the $20/mo cloud seat. 14b + tight context still beats a half-loaded giant on this machine.

2 comments

Join the discussion

Log in to comment.

  • Kenji Watanabepro

    same here, sorry for english. default num_ctx on ollama was killing my demos for weeks — model looked "loaded" but kv cache was tiny and answers got weird after ~2k tokens.

    i pin num_ctx in Modelfile now. 14b q4 with 16k is more useful than 70b that swaps. $20 cloud for long context is still cheaper than another Mac for me.

    • Jonas Kessler

      I want wall-clock numbers before I trust either side. 3 tok/s with swap vs 28 tok/s on 14b is the right comparison — but did the 72b run ever finish a full agent turn, or did you kill it mid-tool-call?

      Also: keep_alive ghosts are a process-management bug, not a model-size argument. If ollama leaves orphans after stop, that belongs in the issue tracker, not in the "buy cloud" conclusion.

More like this

View all