Jev said 0.91 confidence on the wrong payout rail
Tried Jev last week on a transfer-routing decision for our Lagos fintech. Typed output, confidence score, the whole pitch.
It returned rail=swift at 0.91. We almost shipped it. Manual check: that corridor only accepts local NIP for amounts under ₦5m, and the customer was at ₦800k. Confidence looked clean. The type system did not save us from a bad prior.
Two engineers. One almost-bad Friday. Anyone actually gating on the confidence number, or treating it like a vibes meter with prettier JSON?
2 comments
Join the discussion
Log in to comment.
repro-ish: I fed it three corridors with overlapping rules and a stale fee table. Confidence stayed above 0.85 on the wrong one twice. The typed enum looked correct. The world model was outdated.
If your gating is "confidence > 0.8 then auto-route", you're one stale prior away from a chargeback. I want a provenance field on the score, not just a float.
We tried something similar for incident severity triage. Calibrated confidence is nice until the training distribution drifts and the number stays pretty.
Rule we landed on: confidence can bump a ticket priority one notch, never auto-page. A 0.91 that is wrong at 2am is still a pager you own.