vibehacker
Discuss
Chris Vale
11 hours ago

langgraph kept 40 turns and ollama OOMed mid-tool-call

Ollama
Run open models locally and in coding agents

spent friday night wiring a tiny support bot on a mac m2 24gb. qwen2.5-coder:14b via ollama was fine for the first few hops.

then the graph hit ~40 turns with tool results stuffed into state and ollama just died. no clean traceback — just error loading model after the third tool call. activity monitor showed memory pressure yellow for like 10 minutes first.

dropped num_ctx to 8192 and it survived, but then the agent started forgetting the ticket id every other step. anyone else pairing langgraph with local models or am i just being cheap?

2 comments

Join the discussion

Log in to comment.

  • Drift Glyph

    classic. 14b + a fat tool-result buffer is basically asking ollama to duel cursor for the same 24gb. i keep tool payloads on disk and only pass paths + hashes into the graph — ugly, but i stopped seeing error loading model mid-run.

    also check if something else is still holding the previous model. ollama list sometimes lies until you ollama stop the stale one.

  • Linen Pixel

    wait so the ticket id vanishing was just context truncation? that tracks. i had something similar where the agent "fixed" a bug by inventing a new field name once num_ctx got tight.

    do you checkpoint state every N turns or just let langgraph keep the whole history?

More like this

View all