HomeAI NewsAI models flub spatial reasoning puzzles and memorized riddles

AI models flub spatial reasoning puzzles and memorized riddles

Spatial rotation puzzles and memorized riddle variations show where AI still lags human performance.

New tests show that even the best language models struggle with puzzles that most humans can solve, especially visual and spatial reasoning tasks. The same models also tend to stumble when a classic riddle changes in subtle ways, because they lean on memorized answers rather than true understanding.

Puzzles have driven artificial intelligence since its early days, from Arthur Samuel’s checkers program in 1959 to chess and Go benchmarks. Researchers at Columbia University recently found that top models handled only 18 percent of New York Times Connections puzzles in late 2024, then improved to near-perfect scores by early 2025.

For builders, these puzzle results expose specific weaknesses in spatial reasoning and adaptability that differ from human cognition. The same skills that let architects and mechanical engineers rotate objects mentally remain out of reach for language models, and memory can interfere when puzzles vary.

AI puzzle performance is improving quickly, as the jump in Connections scores from late 2024 to early 2025 demonstrates. Yet spatial reasoning and memory adaptability still separate human intelligence from current models, and future benchmark tests will continue to probe those gaps.

What matters

  • Columbia researchers found top models solved just 18 percent of Connections puzzles in late 2024.
  • Spatial rotation tests stun language models despite their ability to process visual input.
  • Google and Illinois researchers showed models over-rely on memorized puzzle answers when variations appear.

Why it matters

Google and Illinois researchers showed models over-rely on memorized puzzle answers when variations appear.

This GenAI News article was prepared in original wording using reporting and materials published by MIT Technology Review AI. Source reference: https://www.technologyreview.com/2026/08/26/1141952/puzzles-ai-models-flub-these-tests/.

Drafted by the GenAI News review pipeline.

latest articles

explore more