New Benchmark Confirms AI Models Still Perform Poorly at Visual Perception
1 min read Image Generation -/5
In short
  • Moonshot AI's PerceptionBench tests how well multimodal AI models can actually 'see,' separate from logical reasoning.
  • No frontier model reaches 60 percent accuracy, with GPT-5.6 Sol leading by a narrow margin.
  • Many supposed reasoning errors occur as early as the image-reading stage.
-/5 (0)
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually 'see,' separate from logical reasoning. No frontier model reaches 60 percent accuracy, with GPT-5.6 Sol leading by a narrow margin. Many supposed reasoning errors occur as early as the image-reading stage. These findings raise questions about the effectiveness of current AI models and highlight the challenges that still need to be addressed to improve visual perception. A final assessment of these models would be premature at this point.