Summary:
- GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.
- GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.
- A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions.
For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted.1
Astra’s results are also a major milestone worth celebrating. From our perspective, Astra represents a noticeable step-function change in frontier model capabilities.
When we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent “proof of achieving AGI.” Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.
I mean, I downvoted because these benchmarks are gamed and only serve as promotion material. Unless there are any other content of high-significance here, I don't believe this post should rank higher on any feed.
Yes, I'm aware many benchmarks are easily gamed and rely mostly on knowledge acquisition rather than true reasoning. I wouldn't post other benchmark results either.
However, ARC-AGI-3 is unique in that it tests systems using dynamic, original game environments where the AI has to learn each game from scratch. Importantly, the score is based on action efficiency, not simply beating each game.
So this benchmark is a bigger signal than most. It should still be judged with skepticism, certainly and doesn't necessarily indicate a fully general intelligence, but it seems like a step-change to me.
I'd encourage you to read more about the benchmark, including the technical paper, if you're interested.
I know about each benchmark, but there are infinite ways to game these benchmarks, even ARC.
How many attempts? How is the final score picked, best, average, worse?
Are tools allowed? Is the model cheating with a tool handcrafted for solving ARC AGI?
Also noticing the score is on the "semi-private" benchmark, if thats even a thing.
The private one has a slower score.
All fair questions. I can imagine OpenAI would cheat in any way they can, but I suppose I'm putting some trust in ARC Prize, or specifically Francois Chollet. I'd be happy to be wrong about this because the implications of it being real are worrying (and I try not to add to the hype/doom). Time will tell I suppose, but I still think it's worth sharing compared to other benchmarks.