vibehacker
News
Hugging Face ·

Prefix cache decides: Qwen3.8-27B NVFP4 serves agents on two RTX 5090s

Scalably’s Sept 19 field report: two RTX 5090s ran Unsloth’s Qwen3.8-27B NVFP4 via vLLM 0.27.0 for 14 days of real agent traffic—82.6% prefix-cache hits, 28k requests, zero engine errors. A faster MoE lost the job on tool recovery (1/20 vs 12/20), so they kept the dense model.

More news

View all

StepFun opens Step 5 Preview API: 600B MoE at $1/$2.70 per M tokens

StepFun opened API access to Step 5 Preview on Sept 20: a proprietary 600B sparse MoE (27B active) with a 1M token context, vision input, and Artificial Analysis pricing of $1/$2.70 per million tokens (Intelligence Index 44). Company reported coding scores still trail GPT 6 Astra and Claude Opus 5 on several benches; open weights are planned for Oct 15…

RuntimeWire

Lawsuit: Anthropic, OpenAI, SpaceXAI, Google illegally agreed to slow AI

A class action filed Friday in Northern District of California claims Anthropic, OpenAI, SpaceXAI, and Google violated antitrust law when their CEOs publicly backed Dario Amodei’s Sept 12 call to pace frontier AI. Named plaintiffs who pay for ChatGPT, Claude, Grok, or Gemini say a coordinated slowdown would cut what subscribers get for their money; the labs had not commented by Saturday…

ABC7

Spotted something we missed? Start a thread.