vibehacker
Discuss

ollama silently capped num_ctx to 2048 and my RAG eval looked smart

Ollama
Run open models locally and in coding agents

Tried Qwen2.5-14B-Instruct-Q5_K_M on my 4090 Friday night. Modelfile had PARAMETER num_ctx 32768. nvidia-smi showed ~11GB so I thought it loaded.

Citation eval hit 0.81. I almost posted it.

Restarted, ran ollama show --modelfile qwen2.5:14b — num_ctx was back to 2048. Long docs got truncated. Real accuracy ~0.34.

sorry english. pin your context or the numbers lie.

3 comments

Join the discussion

Log in to comment.

  • Ash Beacon

    yeah this got me last month. green pytest on truncated context is worse than red.

    i now hit curl -s localhost:11434/api/show -d '{"name":"qwen2.5:14b"}' and assert num_ctx in the fixture before any RAG run. silent defaults are the real bug.

  • Sam Nguyen

    same energy on Mac Studio + LM Studio. UI says 32k loaded, metal quietly drops you when the Q4 doesn't fit. Activity Monitor RAM looks "fine" the whole time.

    I measure tok/s AND the actual context window now or I don't trust the number. 14b with a honest 8k beats a half-loaded giant every time.

  • Elena

    i started printing ollama ps before every RAG bench after this exact trap.

    qwen2.5:14b-instruct-q5_K_M on a 3090 looked "loaded" at 10.8GB. num_ctx was still 2048 until i forced it in the Modelfile AND re-created the model. ollama create from an edited Modelfile ≠ update in place.

    if your citation score jumps 2x after a restart, you didn't improve retrieval — you just stopped truncating.

More like this

View all