has anyone actually got GPT-6 Astra yet or is it just the blog post
openai shipped gpt-6 astra today: https://openai.com/index/gpt-6-astra/
the post says it's rolling out to "a limited set of organizations" first and then plus/pro/api "over the coming days". i'm on pro and i don't have it. codex still gives me sol.
the bits i actually care about from the post:
- in codex it can ask you a question asynchronously and keep working on the parts that don't depend on your answer. if you don't reply it proceeds on the small stuff and waits on the big stuff. that's the feature honestly
- osworld 2.0: 72.6% at ~40 min/task vs sol's 65.7% at ~75 min. so faster more than smarter, for computer use
- they built an eval off the hugging face incident for "does the model go outside its authorized target". sol did that 48% of the time without safeguards, astra 0%. that's the number i'd want to see replicated by someone who isn't openai
so: anyone in the limited set? what's it like on a real repo, not a demo

4 comments
Join the discussion
Log in to comment.
Not in the limited set, sorry. But the ARC Prize post is worth reading before the hype: https://arcprize.org/blog/astra
Same model, two harnesses: 62.7% on ARC-AGI-3 with their standard harness for about $26K, and 99.9% with a "provider adapter" harness that keeps the reasoning state between requests, for about $19K. I think this is maybe the clearest example so far that the harness is half of the score. The model did not change between those two rows.
62.7 vs 99.9 with the same weights. I've been shouting this for a year. The loop is the product. Thanks for the receipt, Kenji.
"$19K for the benchmark run" is the line that got me. Grand that it saturates the thing, but if the cheap harness is the one I can afford, I'm the 62% guy. Nobody's benchmarking the version I'll actually get.
update: still no astra. codex asked me a question mid-task though which it did not do yesterday so maybe the harness update went out first. or i'm imagining it