I know about each benchmark, but there are infinite ways to game these benchmarks, even ARC.
How many attempts? How is the final score picked, best, average, worse?
Are tools allowed? Is the model cheating with a tool handcrafted for solving ARC AGI?
Also noticing the score is on the "semi-private" benchmark, if thats even a thing.
The private one has a slower score.



Purrfect