I find this surprising! I’ve tried this with Opus 4.8 and Fable 5 and they both failed pretty spectacularly. While their vision appears significantly better than Sonnet’s was here, the failures of reasoning remain. To give an example that I found funny/exasperating: it suggested I report that the espresso machine was broken to its owner (that’d be me), failing to realize that the reason it would not turn on is that what Claude told me to press (multiple times!) was not the power button.
I find this surprising! I’ve tried this with Opus 4.8 and Fable 5 and they both failed pretty spectacularly. While their vision appears significantly better than Sonnet’s was here, the failures of reasoning remain. To give an example that I found funny/exasperating: it suggested I report that the espresso machine was broken to its owner (that’d be me), failing to realize that the reason it would not turn on is that what Claude told me to press (multiple times!) was not the power button.
Claudiness = “good at agentic tasks, but bad at vision… and also bad at math”
Claude’s vision is really much worse than ChatGPT, in my experience.