New Benchmark Confirms AI Models Still Perform Poorly at Visual Perception
1 min read
Image Generation
-/5
In short
- Moonshot AI's PerceptionBench tests how well multimodal AI models can actually 'see,' separate from logical reasoning.
- No frontier model reaches 60 percent accuracy, with GPT-5.6 Sol leading by a narrow margin.
- Many supposed reasoning errors occur as early as the image-reading stage.
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually 'see,' separate from logical reasoning. No frontier model reaches 60 percent accuracy, with GPT-5.6 Sol leading by a narrow margin. Many supposed reasoning errors occur as early as the image-reading stage. These findings raise questions about the effectiveness of current AI models and highlight the challenges that still need to be addressed to improve visual perception. A final assessment of these models would be premature at this point.
Source:
-
New benchmark confirms AI models still perform poorly at visual perception — The Decoder (EN-US)