vibehacker
News
Hugging Face / FineEnvs ·

Hugging Face's OpenEnv proxy turns Claude Code, Codex, and OpenCode into RL training environments

An open-source capture proxy sits between an unmodified harness and vLLM, speaks the OpenAI, Anthropic, and Gemini API formats, and records exact token IDs and logprobs so TRL can train on real agent runs; 10 harnesses work today. Training Liquid AI's LFM2.5-2.6B across four harnesses lifted its average solve rate from 42% to 54% (Claude Code 33% to 49%), and the proxy, trainer, tasks, and all seven checkpoints are open.

More news

View all

OpenAI adds opt-in textGrain watermarking to its API and will watermark ChatGPT and Codex text in the EU

Starting Oct 5, API customers anywhere can turn on invisible text watermarking per project or org for select models (off by default), and eligible ChatGPT and Codex users in the EU get it in the coming weeks to meet the EU AI Act. The signal lives in word choice, outputs under 200 tokens and code snippets aren't required to carry it, and the detector is limited to approved researchers for now…

OpenAI

Cursor SDK 1.0.31 lets you steer a running agent mid-turn and swap its system prompt

In @cursor/sdk 1.0.31, run.steer(text) injects a message into the turn in flight (a running foreground subagent moves to the background and keeps going), custom tools can carry MCP annotations like readOnlyHint and destructiveHint , and systemPrompt on Agent.create() replaces Cursor's built in prompt while rules and skills still load. Steering and the prompt swap are TypeScript local agents only, cloud runs return revert to followup , and systemPrompt is enabled per account, so it fails on first send() until yours is on…

AWS ships an aws-ai-ml skill that lets Claude Code, Codex, and Kiro benchmark and size SageMaker endpoints

Install it with npx skills add aws/agent toolkit for aws/skills/aws ai ml (after aws configure agent toolkit , which needs AWS CLI 2.35+ and uv), and your agent can load test a live endpoint, rank instance types for a model from S3, JumpStart, or Hugging Face, and diff two benchmark runs, all as SageMaker Python SDK v3 code you review and run. It runs under your own AWS credentials and asks before driving real traffic, so clean up the endpoints and S3 objects it creates to avoid charges…

NaiveAI open-sources Naive-N0.5-Flash, a 309B MoE coding model with native 1M context under MIT

Built for coding and AI R&D on Xiaomi's MiMo V2.5 Base, the model activates 15.5B parameters per token and reaches 1M tokens of context with sliding window plus DeepSeek Sparse Attention and no full attention layers; weights and inference code are MIT. NaiveAI also plans an API at $0.10 in / $0.40 out per million tokens, but self hosting the full 49 shard checkpoint takes roughly 690 GB of memory…

Together Link routes Claude Code, Codex, OpenCode, and Pi to open models like GLM 5.3 and Kimi K3

One command ( curl fsSL https://link.together.ai/install | bash ) points your existing agent at Together's serverless open models, with an "Auto" router that picks a model once per session (GLM 5.3 vs GLM 5.3 Flash, or Opus 5.5 vs GLM 5.3 if you bring an Anthropic key) so prompt caching keeps working. Together claims over 50% lower spend, shows per session cost next to the Opus 5.5 equivalent, bills to your Together API key, and switching back takes one command…

Together AI Blog

Spotted something we missed? Start a thread.