10 points | by spenvo 19 hours ago ago
5 comments
If its price is almost the same as Opus 5.5 but it’s ultimately not as good, I don’t understand the use case for this model. Perhaps Haiku 5.5 could fill a real need.
It’s better than both fable and Astra? I’m going to start trusting this index less and less.
It's apparently great at Max but extremely costly. Looks best to me at Medium and High. Over that and I'd go Opus. I consider Fable largely obsolete.
the issue is, what will we use? companies are benchmaxxing, so how do you make a bench that cannot be cheated?
Human curated benches aren't accurate enough
The problem with benchmarks is that they are end to end.
You go from one prompt to the final solution, whereas for many of us it is about the experience of iterative, multi turn working.
If its price is almost the same as Opus 5.5 but it’s ultimately not as good, I don’t understand the use case for this model. Perhaps Haiku 5.5 could fill a real need.
It’s better than both fable and Astra? I’m going to start trusting this index less and less.
It's apparently great at Max but extremely costly. Looks best to me at Medium and High. Over that and I'd go Opus. I consider Fable largely obsolete.
the issue is, what will we use? companies are benchmaxxing, so how do you make a bench that cannot be cheated?
Human curated benches aren't accurate enough
The problem with benchmarks is that they are end to end.
You go from one prompt to the final solution, whereas for many of us it is about the experience of iterative, multi turn working.