vibehacker
News
Tuskira ·

Tuskira open-sources self-hosted AI Agent Gateway for LLMs and MCP

Tuskira released an open-source, self-hosted AI Agent Gateway that sits between coding agents (Claude Code, Cursor, Codex) and LLMs/MCP tools, with profile-based access control, call logging, and per-call token cost tracking—no Tuskira account required. Source and Docker deploy live at github.com/Tuskira/tusk-ai-secured-gateway.

More news

View all

Claude Code 2.1.287 ships Mods that rewrite prompts, tools, and UI

Anthropic’s Claude Code 2.1.287 adds Mods—event driven TypeScript/JavaScript hooks (on by default) that can rewrite prompts, tool calls, and UI via the existing plugin system; the /diff pane, agents.md loader, and telemetry now run as mods. Early public mods already include process launching ones, so treat installs like any code with full agent privileges…

Crypto Briefing

Strands Decider 2B: open 2B local decision model for agent workflows

AWS Strands Labs released Strands Decider 2B—an open source 2B parameter decision model (Qwen3.5 2B torso) that picks among fixed options with calibrated confidence in 115ms on a local GPU, for agent routing, tool selection, and guardrails instead of a full LLM call. Weights, training data, and scripts ship on GitHub/Hugging Face with a strands decider CLI and Strands Harness intervention examples…

Strands Agents

OpenAI disrupts Moonshot-linked campaign to extract protected reasoning

OpenAI says it disrupted a July adversarial distillation campaign that tried to extract protected reasoning—16k requests from 4k+ users, related activity across 15k accounts—attributing a core cluster to people associated with Moonshot AI (Kimi). Operators replayed encrypted reasoning across conversations rather than breaking crypto; OpenAI banned accounts, closed the replay path, and says partner hosted deployments still need the same protections…

OpenAI

Magnitude (YC S25): self-optimizing local inference for coding agents

YC S25’s Magnitude (Apache 2.0) is a local inference engine that compiles and tunes kernels on device for agent workloads, claiming up to 2× llama.cpp decode (92% faster on Metal in their benchmarks). A desktop app wires it to Pi, OpenCode, Hermes, Codex, Claude Code, and others with on demand model load/unload…

HN

Spotted something we missed? Start a thread.