vibehacker
News
GitHub / copilot-cli releases ·

Copilot CLI 1.0.92 adds `copilot config` commands and stops passing GITHUB_TOKEN into sandboxed shells

The Oct 5 release adds copilot config subcommands to list, read, set, and remove settings, a Ctrl+E picker before a conversation to choose a local or cloud run, and live stdout/stderr streaming for shell tool calls. Sandboxed shells no longer inherit your ambient GITHUB_TOKEN unless you configure it, so scripts that quietly relied on it will need it set explicitly.

More news

View all

GitHub opens ReviewBench, a public benchmark you can run your own AI code reviewer against

The research preview benchmark scores code review agents on 219 real public PRs across 19 languages, with the full dataset, rubric, and Claude Sonnet 5 judge published, and you can sign in with GitHub, bring a container image plus your own model key, and test on a 25 PR set before a full leaderboard run. GitHub built it to tune Copilot code review, the golden set leans on LLM judging, and leaderboard entries need maintainer approval, so treat the rankings as a starting point…

The GitHub Blog

Superagent open-sources Security-One 27B, a model that scores agent tool calls and inputs as safe or unsafe

Instead of writing out an answer, the Apache 2.0 model returns probabilities for answers you define (like safe and unsafe ) for a prompt, tool call, code change, or alert, so your own code picks a threshold to allow, block, or escalate; Superagent reports catching 599 of 600 BIPIA prompt injections. Prompt injection screening is the best tested use so far, it flagged 12% of tricky benign NotInject inputs, and the numbers are Superagent's own, so tune it on your own traffic and keep real permission checks outside the model…

Superagent Blog

Codex Day 1: GPT-6 Astra and GPT-6.1 Sol run about 50% faster on subscriptions, partners included

The first drop of OpenAI's 28 day improvement or reset run is an inference speedup of roughly 50% for GPT 6 Astra and GPT 6.1 Sol when used through a ChatGPT subscription, including partner tools that use Sign in with ChatGPT like OpenCode, Pi, Amp, and Devin, with nothing to change on your end. API speeds aren't part of it, and one early forum benchmark saw 6.1 Sol get slower mid rollout, so check your own runs…

Spotted something we missed? Start a thread.