ollama json mode keeps inventing camelCase when my Zod is snake_case
hit this three times this week on mistral:small via ollama.
schema says user_id, created_at. model returns userId / createdAt like it read a typescript tutorial and decided my contract was optional. zod .strict() blows up, so the agent "fixes" it by rewriting the schema to match the garbage.
i now coerce keys before parse and treat response_format as a vibe not a contract. anyone else pinning a post-processor for this, or am i just running the wrong model for structured output?
5 comments
Join the discussion
Log in to comment.
what does the raw response look like before you coerce? i've seen models "succeed" with empty objects when the UI just shows a spinner. if you're not logging the bytes, you're debugging vibes.
also: does your agent still rewrite the schema after you coerce, or does that stop it?
logging the raw bytes is the only thing that saved me. once i saw
userIdin the payload while the UI showed a green "structured" badge i stopped trusting the badge.repair step before zod > swapping models every tuesday. i keep a sticky of known remaps per model id.
same class of bug as agents discovering .gitignore. output that passes "looks like json" is still untrusted input.
i pipe through a 12-line jq remap (snake_case only) before zod. costs nothing, saves the friday rewrite loop. mistral:small is fine for chat; for contracts i bumped to qwen2.5 and it still drifts ~1/20. post-processor stays.
same on my M2. qwen2.5:14b drifts less than mistral:small but still invents createdAt when the schema says created_at about once a night.
i pin num_ctx and treat the json mode output as untrusted. the jq remap is boring and it works. still rebooting after silent ollama deaths though
yeah mistral:small does this constantly for me. i stopped trusting response_format after it renamed half my invoice fields to camelCase and the agent "helpfully" updated the zod schema to match.
now i strip keys through a tiny remap and .strict() after. costs like 8 lines. still blows up ~1/15 on llama3.2:3b though — are you pinning a larger model for the parse step or same small one for everything?