ollama silently capped num_ctx to 2048 and my RAG eval looked smart
Tried Qwen2.5-14B-Instruct-Q5_K_M on my 4090 Friday night. Modelfile had PARAMETER num_ctx 32768. nvidia-smi showed ~11GB so I thought it loaded.
Citation eval hit 0.81. I almost posted it.
Restarted, ran ollama show --modelfile qwen2.5:14b — num_ctx was back to 2048. Long docs got truncated. Real accuracy ~0.34.
sorry english. pin your context or the numbers lie.


3 comments
Join the discussion
Log in to comment.
yeah this got me last month. green pytest on truncated context is worse than red.
i now hit
curl -s localhost:11434/api/show -d '{"name":"qwen2.5:14b"}'and assertnum_ctxin the fixture before any RAG run. silent defaults are the real bug.same energy on Mac Studio + LM Studio. UI says 32k loaded, metal quietly drops you when the Q4 doesn't fit. Activity Monitor RAM looks "fine" the whole time.
I measure tok/s AND the actual context window now or I don't trust the number. 14b with a honest 8k beats a half-loaded giant every time.
i started printing
ollama psbefore every RAG bench after this exact trap.qwen2.5:14b-instruct-q5_K_M on a 3090 looked "loaded" at 10.8GB. num_ctx was still 2048 until i forced it in the Modelfile AND re-created the model.
ollama createfrom an edited Modelfile ≠ update in place.if your citation score jumps 2x after a restart, you didn't improve retrieval — you just stopped truncating.