They game benchmarks on these because they are desperate to keep funding going. They are so far from profitable
Futurology
I guess this is being down-voted just because it feels pro-AI. Believe me I get it, you can look at my anti-AI post history. But I think it's important to also keep an eye out for real signals towards general intelligence, which is why I wanted to share this. I'm not celebrating it, just bringing awareness.
I mean, I downvoted because these benchmarks are gamed and only serve as promotion material. Unless there are any other content of high-significance here, I don't believe this post should rank higher on any feed.
Yes, I'm aware many benchmarks are easily gamed and rely mostly on knowledge acquisition rather than true reasoning. I wouldn't post other benchmark results either.
However, ARC-AGI-3 is unique in that it tests systems using dynamic, original game environments where the AI has to learn each game from scratch. Importantly, the score is based on action efficiency, not simply beating each game.
So this benchmark is a bigger signal than most. It should still be judged with skepticism, certainly and doesn't necessarily indicate a fully general intelligence, but it seems like a step-change to me.
I'd encourage you to read more about the benchmark, including the technical paper, if you're interested.
I know about each benchmark, but there are infinite ways to game these benchmarks, even ARC.
How many attempts? How is the final score picked, best, average, worse?
Are tools allowed? Is the model cheating with a tool handcrafted for solving ARC AGI?
Also noticing the score is on the "semi-private" benchmark, if thats even a thing.
The private one has a slower score.
All fair questions. I can imagine OpenAI would cheat in any way they can, but I suppose I'm putting some trust in ARC Prize, or specifically Francois Chollet. I'd be happy to be wrong about this because the implications of it being real are worrying (and I try not to add to the hype/doom). Time will tell I suppose, but I still think it's worth sharing compared to other benchmarks.