The ARC-AGI benchmark has long been regarded as a tough measure of AI's "general reasoning ability"—it doesn't rely on rote memorization or pattern matching, but purely tests whether a model can adapt flexibly to abstract tasks it has never seen before.
GPT-5.2 scored 72% on ARC-AGI-1 and 18% on ARC-AGI-2.
What Do These Numbers Mean?
The 72% on ARC-AGI-1 already surpasses the vast majority of models. But ARC-AGI-2 is the real touchstone—its difficulty is several notches higher. The 18% shows that current LLMs still have very limited capability when facing novel, untrained abstract reasoning tasks.
To draw an analogy: it's like a student who scores high on exams but panics when faced with question types they've never seen before. This gap precisely reveals the chasm between AI that "looks smart" and truly general intelligence.
OpenAI's Strategic Shift
A clear trend with GPT-5.2 is cost optimization. Compared to the previous generation, the cost of calling the same tasks has dropped significantly, while reasoning ability has been pushed up a notch.
OpenAI is evidently pursuing two paths:
- High-end line: competing on absolute capability ceiling
- Economy line: making existing capabilities cheaper
This strategy forms an interesting contrast with DeepSeek's approach of using MoE to compress costs. One is lowering prices from above, the other is raising capabilities from below—the point where they meet may be the industry's equilibrium price.
For those tracking AGI progress, the 18% on ARC-AGI-2 is a more informative number than various MMLU or HumanEval scores—because it directly measures generalization ability, not memory.
Sources: CocoLoop, ARC Prize official blog, OpenAI technical report