actions/cache restored a dead .venv and my agent spent 40 min fixing nothing
hit this on a macos-14 runner tuesday night.
poetry.lock was fine. cache key used hashFiles('**/poetry.lock') so it restored a .venv from last week with pydantic 2.8 while the lock said 2.10. locally pytest green. in CI: ImportError: cannot import name 'model_validator'.
claude code kept rewriting my conftest. never touched the cache key. i nuked the actions/cache@v4 restore and the "bug" vanished in one run.
anyone else hashing lockfiles but caching the whole venv?
5 comments
Join the discussion
Log in to comment.
same class of lie. promptfoo was green because it loaded fixtures from
node_modules/.cachethat still had the old system prompt string.i now bust cache on any change under
prompts/and print the resolved prompt hash in the job log. if the agent can't see that hash, i don't let it "fix" eval failures.prompt hash in the log is good. we also fail the job if CACHE_HIT=true and lockfile mtime is newer than the cache key timestamp. agents love to "fix" import errors by editing the wrong file.
had this with n8n on a self-hosted runner. cache hit at 23:11 UTC, agent looped on ImportError until 01:40. bill was ugly.
i started dumping cache hit/miss into the job summary so i stop trusting the green check. hashing the lockfile is necessary but not sufficient if the venv itself is what you restore.
lost a thursday night to this before a soft launch. two engineers, one runner, dead venv from tuesday.
now the agent gets one rule: do not touch tests or conftest until you print whether the cache restored. still breaks sometimes but at least it asks first.
I hash poetry.lock AND the python version string. Still not enough if you restore whole .venv.
On our MCP CI we switched to
pip install --targetinto a path we never cache. Slower by ~90s. Agent stopped inventing ImportError fixes. Worth it.