

ARC-AGI-3
#27 v Reasoning modelyUnknown · v3 · od 25. März 2026 · 2× · naposledy 02. 10. 2026
18
Momentum
ARC-AGI-3 is an interactive reasoning benchmark by the ARC Prize Foundation that tests AI agents in novel, instruction-free game environments. Instead of static puzzles, agents must explore rules, infer goals, build world models, and learn continuously across levels without guidance. The benchmark comprises hundreds of handcrafted turn-based environments; humans score 100% while frontier AI scored below 1% at launch. Access is via an open REST API and Python toolkit (MIT license), with results published on a public leaderboard.
Vývoj momenta
04.07.02.10.
Vlastnosti
| Key Benchmark (%) | Best score (leaderboard, Sep 2026): GPT-6 Astra 62.7% (up to 99.9% depending on harness); humans 100%; frontier AI at launch 0.51% |
| License | ARC-AGI Toolkit/benchmarking code: MIT license; competition submissions must be CC0/MIT-0; private test set restricted |
| Multimodality | Visual/interactive: 2D pixel-grid game environments (e.g., 64x64), no language instructions, agents act via real-time actions |
| Platform | Web-based game at arcprize.org/tasks; agent access via REST API and Python toolkit (arcprize/ARC-AGI on GitHub) |
| Release Date | March 25, 2026 (official launch) |