vibehacker
News
Reflection Blog ·

Reflection previews Beam, a 501B open-weight MoE coding model with Apache 2.0 weights due this month

Beam activates 23B of its 501B parameters per token, has 1M-token context, and Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal Bench 2.1, roughly level with GLM 5.2 at 3–4x less inference compute but behind Kimi K3 and GLM 5.3. It's waitlist-only for now, with weights, a technical report, and fine-tuning tooling promised later in October.

More news

View all

Superagent open-sources Security-One 27B, a model that scores agent tool calls and inputs as safe or unsafe

Instead of writing out an answer, the Apache 2.0 model returns probabilities for answers you define (like safe and unsafe ) for a prompt, tool call, code change, or alert, so your own code picks a threshold to allow, block, or escalate; Superagent reports catching 599 of 600 BIPIA prompt injections. Prompt injection screening is the best tested use so far, it flagged 12% of tricky benign NotInject inputs, and the numbers are Superagent's own, so tune it on your own traffic and keep real permission checks outside the model…

Superagent Blog

Codex Day 1: GPT-6 Astra and GPT-6.1 Sol run about 50% faster on subscriptions, partners included

The first drop of OpenAI's 28 day improvement or reset run is an inference speedup of roughly 50% for GPT 6 Astra and GPT 6.1 Sol when used through a ChatGPT subscription, including partner tools that use Sign in with ChatGPT like OpenCode, Pi, Amp, and Devin, with nothing to change on your end. API speeds aren't part of it, and one early forum benchmark saw 6.1 Sol get slower mid rollout, so check your own runs…

CodeAF is an open-source Go coding harness built for open models like DeepSeek, Qwen, and GLM

AgentField's Apache 2.0 CodeAF ships as one 53 MB binary ( curl fsSL https://agentfield.ai/get/codeaf | bash ) that turns each request into its own task on its own branch, runs it on a per task crew of open models via OpenRouter, provider keys, or local Ollama, and claims first place among ten harnesses on DeepSWE with DeepSeek V4 Flash. It's an early preview with anonymous telemetry on by default, which CODEAF TELEMETRY=off or DO NOT TRACK=1 turns off…

Spotted something we missed? Start a thread.