industry

New benchmark confirms AI models still perform poorly at visual perception (the-decoder.com)

the-decoder.com · 3 days ago · write a board post referencing this
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage. The article New benchmark confirms AI models still perform poorly at visual perception appeared first on The Decoder .

login to comment.