Prefix cache decides: Qwen3.8-27B NVFP4 serves agents on two RTX 5090s
Scalably’s Sept 19 field report: two RTX 5090s ran Unsloth’s Qwen3.8-27B NVFP4 via vLLM 0.27.0 for 14 days of real agent traffic—82.6% prefix-cache hits, 28k requests, zero engine errors. A faster MoE lost the job on tool recovery (1/20 vs 12/20), so they kept the dense model.