vibehacker
News
PyPI / Rio ·

Rio 0.7.1: one-shot coding agent that uses a Jupyter notebook as its context

Ju Lin's MIT-licensed Rio runs a task to completion with no chat or mid-run steering, acting by patching Python, shell, and %%edit cells in a live Jupyter kernel instead of growing a conversation history. Install with uv tool install rio, log in to Anthropic, Codex, or GitHub Copilot, then rio run "Fix gh issue 123." or point it at a task.md.

More news

View all

Microsoft's Agensh: leaderless coding-agent teams keep improving from 1 to 1,024 agents

Agensh drops the central orchestrator: each worker claims its own sub task, builds and tests it, and merges into a shared Git repo while posting findings to a shared board, all on top of a Copilot single agent harness with a 6 hour budget. On ProgramBench's five hardest tasks with GPT 5.6 sol, going from 1 to 128 agents lifted the mean test pass rate from 19.31% to 28.78%, and 1,024 agents pushed pandoc from 33.89% to 55.06%, though the paper doesn't report what large teams cost…

Strata runs the 125B Qwen3.8-Flash-Next on a 12 GB gaming GPU, fully local

MIT licensed Strata is a one click Windows/Linux installer that serves Qwen3.8 Flash Next behind OpenAI and Anthropic compatible APIs on localhost, so you can point Claude Code, Cursor, or Codex at it; the author measured 53–94 tokens/s on an RTX 5070 (12 GB) with 64 GB RAM. You need 32 GB+ RAM and about 80 GB of disk, and the PC can freeze for 1–3 minutes while the model loads…

Show HN: Offrun, a free Mac workspace for running Claude Code, Codex, and other agent CLIs

Offrun runs the agent CLIs you already use (Claude Code, Codex, AGY, Grok Build) side by side on your own accounts, giving each agent its own git worktree, letting a second agent review diffs before commit, and moving a chat to another login when one hits its limit. It's free, needs an Apple Silicon Mac on macOS 14+, and says prompts go straight to your provider, never through its servers…

Show HN / Offrun

Benzi: compiler-backed coding agent hits 78.2% on SWE-bench Verified for $37

Benzi parses your repo with tree sitter into a resolved map of symbols, call edges, data flow, and inheritance, then has the model query it through 36 structured tools instead of reading raw files; it ships as a VS Code extension, a pip install benzi package, and an MCP server. Its authors report 391/500 on SWE bench Verified with DeepSeek v4 flash at pass@1 for $37.33 total, a self run result that includes a review pass, not just the index…

GitHub / Benzi

Show HN: Loa, a local-first coding agent that plans work as a mutable DAG

Loa runs long coding tasks against your own OpenAI compatible endpoint (Ollama works), building acceptance criteria and a replannable DAG, then rebuilding context from state each step instead of compressing chat history; the author suggests 27B–32B local models. It makes many model calls by design, so skip metered cloud APIs, and use the optional Docker sandbox since host mode runs shell commands with your privileges…

Show HN / GitHub

Spotted something we missed? Start a thread.