The unwritten benchmark: why multimodal AI fails at abstract perceptual reasoning
New research shows that top AI models like GPT-4o fail at abstract perceptual reasoning, scoring under 10% on the new 'Unwritten Benchmark' test.
Current multimodal AI models struggle with abstract perceptual reasoning, failing to infer information from dynamic physical processes like handwriting. While humans excel at this due to internal mental models of physics, AI models currently lack the ability to synthesize complementary sensory inputs, highlighting a critical limitation in their capacity for real-world causal reasoning.
New research shows that top AI models like GPT-4o fail at abstract perceptual reasoning, scoring under 10% on the new 'Unwritten Benchmark' test.
The Unwritten Benchmark is a research challenge designed to test if AI models can infer words being written by observing only the sound of the pen and the movement of the hand, without seeing the actual ink or the final written result.
AI models perform poorly because they rely on pattern matching rather than simulating physical processes. They lack an intuitive grasp of how the kinematics of writing relate to the resulting symbols, causing them to fail when the visual 'answer' (the ink) is removed.
The paradoxical fusion effect occurs when an AI model performs worse when given both video and audio data compared to when it is given only one of those modalities. It suggests the model cannot effectively integrate, or 'fuse,' complementary sensory data to solve a problem.
It means current models are not 'reasoning' in the human sense. They are highly efficient at statistical pattern recognition on static data, but they lack the generative mental models required to understand physical causality in dynamic, real-world environments.