61 comments

  • aliljet a day ago ago

    It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...

    • nl 20 hours ago ago

      This is completely the wrong read!

      Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task.

      Additionally, if you are using OpenAI you have the option to pay a bit more and get Astra which - despite the benchmarks - does outperform Sol on some things.

      Also, people are - rightly - very wary of Google's benchmaxxing tendencies. I think lots of people remember Gemini 3.0 (I think?) which benchmarked amazingly, but as soon as you used it would go off-track and needed constant babysitting if you wanted to use it for agentic work.

      • re-thc 12 hours ago ago

        > Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task.

        Via the API. The $200 OpenAI plan just got cut and most say general quotas got cut before that so for users on a plan the numbers might be different.

        • nl 11 hours ago ago

          Yes well obviously it's comparing API vs API and subscriptions are much cheaper on both.

    • scrollop 15 hours ago ago
      • ozgung 13 hours ago ago

        I wonder how they do this nerfing thing. One candidate is slightly decreasing the number of chain of thought tokens for each effort level. It must be something they do to meet increasing demand. Also the significant drop before a new release is because of reallocating the resources.

        Methodology section in some of these benchmarks doesn’t say if they use subscription or API. API usage may not be nerfed as much as subscription.

        They must be using the same lever to “pace the frontier”. All of the best effort models from different companies have similar scores. There is no standard definition of “max” effort level.

    • tomrod a day ago ago

      They own their hardware. That vertical integration alone probably saves oodles because they can reconfigure to their needs as opposed to individually negotiating data centre plans. I can't imagine the complexity both OpenAI and Anthropic have to maintain for their deployments.

      • mkotlikov 21 hours ago ago

        It's still not enough though, they're buying compute from SpaceX's colossus data centers.

        • genxy 19 hours ago ago

          That is google trying to keep their 10% worth something.

          • largbae 19 hours ago ago

            And it totally worked too.

        • tomrod 20 hours ago ago

          They may be buying it, but is there any indication they are using it for consumer-facing stuff? I would suspect given the scale there are many different things. I wish I knew more about standardization for cloud data center workloads though

        • notatoad 20 hours ago ago

          i don't know if you could say they're buying it "happily". more like begrudgingly.

          the deal has a 30-day cancellation policy, and they raised a bunch of debt around the same time to fund their own datacenter expansion.

          • fragmede 17 hours ago ago

            and they shut down Stadia so they'd have some extra GPUs to use for it as well!

    • sroussey 20 hours ago ago

      GPT-6.1-sol costs less than half of Gemini and way less than Anthropic on the cost per task chart of the listed parent page.

      • Melatonic 12 hours ago ago

        Suspect they will release Gemini 4 flash not too long from now as the cheaper option

    • aleqs a day ago ago

      > Opus 5.5 right now because it's nearly unlimited use

      Anthropic has some of the lowest usage per $ in general, not sure what you're taking about.

      • nl 20 hours ago ago

        This is no longer the case.

        Opus and Sol usage levels vs the API are currently roughly the same, but Opus 5.5 outperforms at low and medium effort levels.

      • jjice a day ago ago

        I don't think the OP and your comments are mutually exclusive. If Anthropic usage is lower, and the OP considers it basically unlimited for their case, that just means that they would have virtually unlimited usage with other plans.

        • aleqs a day ago ago

          Nothing about it is unlimited or even close to it. Just because you only use 1gb of your 2gb data plan, doesn't mean you have unlimited data.

          • 8n4vidtmkvmk 16 hours ago ago

            What if you use 100GB of your 10TB plan? Is it virtually unlimited then?

            It's all relative. Some people just can't use up their quotas with their normal usage.

            • aleqs 15 hours ago ago

              That just means your use is limited not that the plan is unlimited.

      • UltraSane 21 hours ago ago

        Opus 5.5. on medium effort is very good a writing code and provides a lot of tokens per 5 hour session. It is a fantastic value for $20/month.

    • zozbot234 a day ago ago

      It's not as smart as Claude Opus 5.5 High according to the AA benchmark. Looks like a big fat nothingburger so far, though it's possible that future fine-tuned checkpoints of the same pretrained model will do a lot better.

      • mpyne a day ago ago

        > It's not as smart as Claude Opus 5.5 High according to the AA benchmark.

        If it's smart enough to do the job then it won't matter that Opus is smarter. At the right price and performance, at least.

        • jug 10 hours ago ago

          I checked price per AA task and yeah Argon vs Opus 5.5 High are practically neck and neck both on "intelligence" and cost per task. So if you aren't pushing higher on Opus (which I think can quickly become very expensive and I always treat Max and those like "benchmark settings"), I think this becomes more of a matter of which platform you like more or prefer for various reasons.

          Honestly, I think this will be a trend in 2027 when all those models become "good enough" for elite coding and whatever. I predict they'll have to branch out more and build their platform to differentiate themselves from each other, maybe even in terms of branding and trust, marketing towards various demographies, youth vs elderly, students vs employees, etc.

        • petesergeant 13 hours ago ago

          While that’s true, I have found Gemini models to be exclusively good for data extraction, and absolutely terrible at everything else.

          I pay for lots of models because they’re good at different things: $20 a month each for Grok and GLM have easily paid for themselves by finding bugs that my main work models didn’t, but I’m yet to have any Gemini model find a real bug, and Gemini’s results for general work will sometimes border malicious compliance, when it’s not having a hissy fit about some imagined issue.

  • mlmonkey 20 hours ago ago

    Meanwhile I'm still being offered Gemini 3.1Pro on gemini.google.com :-D

    https://imgur.com/a/h96yg5t

    This is on a $20/mo paid plan :cry:

    • brainwad 15 hours ago ago

      That is considered a consumer-level surface so the prompt and features layered on top are more important than the model underneath. If you want the latest models you should use Antigravity, it's the "pro" interface.

    • radicality 18 hours ago ago

      On my paid personal Workspace account I’m even more behind. I see 3.1 Pro, 3.6 Flash (New). Can’t not laugh at that ‘New’.

  • algoth1 a day ago ago

    The most impressive jump for me is in the low hallucination rate, which is specially impressive given how bad Gemini current models are on this regard

  • ai-x a day ago ago

    Note: Google can sell their tokens at cost if they really want to drive out competition, but long term they are better off by everyone making a healthy margin (and Google does a double-dip by also selling compute, services).

    So, like any optimal game theory move, they are better off not starting a price war

    • gradus_ad 21 hours ago ago

      Important to remember AI is a direct assault on their actually profitable business: search and ads.

      OpenAI and Anthropic are existential threats to Google and it will operate accordingly.

    • pooper a day ago ago

      A slightly different question would be can they afford to "sit out" of any price war?

      • ai-x 20 hours ago ago

        As long as they are serving the models through Google Cloud and have a SOTA model for

        a) internal use b) embedding in their products c) stay abreast with capabilities

        they can sit it out.

    • miohtama a day ago ago

      Google is losing in the market share - their share is 0. Anthropic revenue was $60B in the last 12 months. Google is not growing the pie, it needs to buy its way to the market share and to a seat in the table.

  • mchusma 18 hours ago ago

    I’m guessing this is considered something like a C grade from Google if they are being honest with themselves.

    After being nowhere near the frontier for a long time, they are pre announcing a model that ranks 3rd, roughly on par with models today that are cheaper.

    Good for them to think about releasing to stay in the frontier game.

    (I do think 3.7 flash was a solid release, so they are around the conversation. And their image and audio and live models are good)

  • yipinwong a day ago ago

    I am still not sold on Gemini 4 Argon yet from the chart.

    The price is enticing for cost per tasks, but let's see how it goes.

    I have montly (cheapy) sub to gemini models and has been underwelming and lowered the tier.

  • lhk931122 19 hours ago ago

    I'm not sure Google can cut in when Claude and ChatGPT already got. I'm using both, but I'll keep track of whether Google can make it good enough for me to use also this Gemini, or cancel one of the two (Claude and ChaGPT) for it.

    • epolanski 13 hours ago ago

      Companies out there are on google cloud or microsoft offerings and getting Gemini in their bundle, they aren't going through lawyers, etc, to provision from Anthropic or OpenAI just because they look a bit better on nerd benchmarks.

  • godbox a day ago ago

    What a snooze fest. Another model that does not meaningfully improve on intelligence or price compared to its peers. Google has basically announced that they've "caught up" with the rest. I think they've been doing great work in the Flash department so seeing this is... underwhelming?

    • losvedir a day ago ago

      The only models beating it are on "max", while this is "high". There's no guarantee that those effort / reasoning levels compare, but there's almost certainly an "xhigh" or "max" version later which will score higher there.

    • augment_me a day ago ago

      I found that Google does not benchmaxx as much as the other providers. Of you look at real-case evaluation like lm-arena, even the 3.8 flash is often near the top despite its benchmark index being worse.

      • sourweasel 21 hours ago ago

        I noticed this recently with the user submitted benchmarks on Kaggle. For these obscure tests that the models haven't seen, Gemini 3.8 is often on par or beating other frontier models. Gemini is a bit shite at the game checkers though, for some weird reason it performs poorly on those benchmarks.

    • piyh a day ago ago

      > does not meaningfully improve on intelligence or price compared to its peers

      $10 per million output tokens isn't improving on frontier price?

      • godbox 16 hours ago ago

        That is a promotional price. The regular price is double that.

    • netdur 21 hours ago ago

      flash on ai studio is my fav model to chat with, by miles

    • dzhiurgis 21 hours ago ago

      It's half cost of gpt for same intelligence. Nearly 4x cheaper than claude.

    • enraged_camel a day ago ago

      Based on some... rumors I've heard, this is their "Pro" offering. There is supposed to be an Ultra coming as well.

  • jwpapi a day ago ago

    Always the last model that gets announced is the best. The labs always know in advance.

  • small_model 20 hours ago ago

    5 points off Opus 5.5 on AA, not a good release. Falling behind and not able to catchup. Ant probably has opus 6 in the works. Fumbled so hard on this, they should have owned AI.

  • dang a day ago ago

    Related ongoing thread:

    Gemini 4 Argon - https://news.ycombinator.com/item?id=49913571

  • dom96 a day ago ago

    I'd love to run it on my benchmark but alas, Google not making it public prevents this.

  • anuragdaram a day ago ago

    Will have to check what the pricing would be for this model.

    • krat0sprakhar a day ago ago

      > Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.

      https://blog.google/innovation-and-ai/models-and-research/ge...

      1/5th the price of Astra and Fable

      • sroussey 20 hours ago ago

        Same price as GPT-6.1-sol however.

        • anuragdaram 3 hours ago ago

          That would be great if it can be as good as those models. I was very excited about 3.8 Flash but it doom looped while executing a coding task and I stopped using it then. And will check how the token usage is like for this one.

  • A_D_E_P_T a day ago ago

    Trading punches in the benchmarks with Mimo v2.6 and 6.1-Sol (both very cheap!), and decidedly inferior to Opus 5.5. I'm afraid this looks unimpressive. Rather comical that they're delaying its launch "for safety reasons".

    • asdfasgasdgasdg a day ago ago

      From the charts, it's similar in intelligence and cost per task to both Opus 5.5 (high) and 6-Astra (max). It would be better if it were more intelligent and less expensive, but I don't see a reason to expect it to have better performance than models released around the same time.

    • godbox a day ago ago

      To play the Devil's advocate, Claude loves chugging tokens, while Gemini appears to be quite a bit more conservative and efficient. I believe AA's price per task breakdown reflects this.

    • thereitgoes456 a day ago ago

      What are you talking about? It’s comparable to Opus 5.5 on “high” (54 vs 53; $1.82 vs $1.99), crushes every model except the most modern OAI/Ant ones, has way lower hallucination than every existing model and probably broader support for multimodal like existing Gemini models. This is so ludicrously off base.