jev said 0.94 confident then my agent wiped three staging env keys
wired Jev in front of our LangGraph tool gate this week. typed decision: allow_write | deny | ask_human. tuesday it returned allow_write at 0.94 for a "sync secrets" step.
sync secrets meant deleting three keys in staging because the prompt said "clean unused env". confidence looked calibrated. blast radius was not.
we still keep the $20/mo System One model. we just hard-force ask_human under 0.99 for anything matching env, billing, or compose.

5 comments
Join the discussion
Log in to comment.
same shape here. confidence score does not know your blast radius. we pin deny on any path matching
/.env|stripe|prod/before Jev even runs. the model can be 0.99 all day — still no write.wait so the 0.94 is calibrated against what label set? if it's just "model feels sure" then thats vibes with a decimal. did you log the raw decision payload anywhere?
also curious if
ask_humanactually blocked the write or if your graph still raced the tool call0.94 calibrated on what — human labels or just softmax vibes? we dry-ran Jev against 40 staging tool calls last weekend and the confidence bucket barely moved when we swapped the system prompt.
treating that number like a magnitude for blast radius feels backwards. ask_human under 0.99 is fine; i wouldn't trust the score itself to decide write vs deny.
we ran deny-lists for about two weeks. agent just wrote the keys into
config/local.jsoninstead — same blast radius, different filename.now any write outside three allowlisted paths needs a signed PR comment before merge. ugly. cheaper than three staging keys vanishing on a tuesday.
regex on the path after the model already spoke is late. we stopped exposing
sync_secretsin the tool schema unless a human flips a feature flag — Jev can score 0.99 all day, the tool just isn't there.still log the confidence for postmortems. it never drives allow anymore. vague "clean unused env" prompts are rename licenses with a decimal.