ollama idled at 40W overnight and still didn't finish my eval
left qwen2.5:14b running an eval harness on a Mac Mini M2 friday night. Activity Monitor said ~40W the whole time. saturday morning: 312/500 cases done, context still silently clipped somewhere around 8k despite me setting num_ctx.
i redid the last 188 on a $20/mo cloud endpoint in under an hour. tokens-per-watt looked cute until the wall clock said otherwise.
anyone actually finishing offline evals on apple silicon without babysitting the context window?
5 comments
Join the discussion
Log in to comment.
same energy. i left mlx + a 7b looping in tmux, blamed the cafe wifi for three hours, then checked
nvidia-smion the wrong machine.berlin tip: if the fan is quiet and the progress bar crawls, you are not offline-heroing, you are napping with a gpu. cloud for the long eval, local for the smoke tests.
honest question: how many hours of babysitting before the $20/mo endpoint is cheaper than your sleep?
i cancelled two "local forever" setups this year. one weekend of unfinished evals already costs more than three months of cloud. keep ollama for smoke tests, pay for the overnight batch.
curious what your harness looks like — are you pinning
num_ctxin Modelfile, or only atollama run?i keep a Notion table of these fails. half the time the model tag silently resets context on reload. a one-line AGENTS.md that says "print
num_ctxbefore every batch" has saved me more than swapping models.sorry english. same clip with qwen2.5:14b on M2 Pro last month.
ollama showsaid 32768 but the harness logs only saw ~8k before repeats. ModelfilePARAMETER num_ctx 16384only sticks if youollama createthe tag again — not justrun -c.i also print num_ctx before every batch now. saved more nights than swapping models.
40W overnight is not a success metric. if the harness cannot print live context length every N cases, you are guessing.
Activity Monitor lied to me for years on agent jobs too. instrument the runner, not the fan curve.