Groq nuked my free-tier eval at case 91
Ran a 180-case prompt eval against llama-3.3-70b on Groq this morning. First 90 finished in under 4 minutes which felt illegal.
Then the free tier hard-stopped me mid-batch. Error was just rate_limit_exceeded with a retry-after of like 14 hours. Switched the remaining cases to OpenRouter and watched the meter jump ~$11 before lunch.
Anyone else keeping a cold cache of fixtures for this, or are we all just eating surprise bills?
5 comments
Join the discussion
Log in to comment.
yeah this bit me last week. i started treating free tier like a demo lane only — anything over ~40 calls goes straight to a paid key with a hard $5 daily cap in the wrapper.
still cheaper than me sitting there clicking retry for an hour.
wait — how are you wiring the $5 daily cap? env flag in the wrapper, or something that actually cuts the key mid-run?
i tried a soft counter once and it still blew past $8 while i was in a meeting. want the dumb version that just hard-fails.
Cache the fixtures. I batch at 16, dump request/response pairs to disk on first green run, then replay locally for regressions.
Live Groq only for the weekly gold set now. Saved me both rate limits and the "why did case 47 flip overnight" panic.
Fixtures help until the model drifts and your "green" suite is lying to you.
I keep a 20-case live canary on a paid Groq key with a hard $3/day stop in the client. Rest of the suite is cached. When canary flips, I regenerate fixtures, not the other way around.
That OpenRouter $11 jump is exactly why I stopped treating free tier like production.
Same movie on a meetup demo last month. llama-3.1-8b on free Groq looked snappy for the first 30 requests, then
rate_limit_exceededon stage while people were watching.Now staging always has a Fireworks fallback keyed off HTTP 429. Boring, but the room never sees the spinner of shame again.