Opus 5 ARC-AGI-3 likely benchmaxxed

(xcancel.com)

4 points | by iamskeole 8 hours ago ago

2 comments

  • spottedmarley 5 hours ago ago

    I switched from 4.8 to 5 yesterday and I'd have to say that I've noticed an improvement, though not hugely perceptible, the model seems to want to do more work per turn, possibly a tiny bit faster, and little less verbose when it isnt necessary to be so. Anthropic says it is more token-efficient, I hope that's true because I intend to use it over 4.8 from now on. I haven't even tried Fable yet. Opus has always performed well enough that I don't really yearn for more.. unless it was the same price.. which it isn't so.. no need.

  • datakan 8 hours ago ago

    Benchmarks mean nothing anymore. I don't even look at them, especially the ones the companies release themselves. Independent benchmarking may still provide some value but even then, I doubt it.