ollama qwen2.5:14b claimed 32k context then silently ran at 8192
spent tuesday night in the garage chasing why my rag answers went stupid after page 3.
ollama show qwen2.5:14b-instruct says 32k. Activity Monitor shows ~28GB resident on the Mac Studio. ask it to summarize a 12k-token dump and it starts inventing section headers that aren't in the file.
checked /api/generate — num_ctx was still 8192. nowhere in the UI screams that. flipped it to 16384, answers got boring again (good). dropped back to 14b + tight chunks instead of chasing a half-loaded 70B.
anyone else treating the model card context as marketing until they print the actual request?


5 comments
Join the discussion
Log in to comment.
yes. default num_ctx quietly kills demos. i pin it in Modelfile now, otherwise every restart lies.
sorry english — also check
prompt_eval_countin the response. if it stays under ~8k while you feed 12k, the card number is marketing.repro that worked for me: dump the raw
/api/generatebody before you read the answer. ifoptions.num_ctxis missing, ollama quietly uses 8192 even when the card says 32k.i also assert
prompt_eval_countagainst expected tokens in a tiny pytest. green badge that skips that check is useless.same on a 4090 with vLLM. people paste "128k support" from the HF card and never look at
--max-model-len.i log
n_ctxon every server boot now. if the number is wrong the box refuses to start. painful once, saves a weekend.yeah. on the m2 i stopped trusting
ollama showafter it said 32k and metal still allocated like an 8k window.i keep a sticky of working num_ctx per tag now. qwen2.5:14b-instruct only behaves once i force 16384 in the Modelfile and pass it again in the generate call. one without the other and it lies again after a reboot.
does the ollama app surface this anywhere, or do you only catch it when answers get weird?
i only noticed after Activity Monitor looked fine and summaries started inventing section 4. curious if LM Studio at least yells on the clamp or if it does the same silent thing.