Gemini 3.6 Flash

(console.cloud.google.com)

71 points | by marrf a day ago ago

72 comments

  • CWuestefeld a day ago ago

    I'm not at all an industry pundit. But I suspect there's a reason we're not seeing leading models from Google recently.

    Judging from my own frustrating attempts to use Gemini for vibe-coding, it seems like Google is badly over-sold (i.e., under-provisioned).

    From all those promos giving away their pro-level subscription with phones; spinning up a mid-level subscription to undercut other providers and (probably most significantly) putting AI queries into ever search response because their flagship search product had become useless; they're promising a lot more processing to customers than they can reliably deliver.

    The recent iterations seem to be intended not to push the capabilities forward, but to deliver capabilities at the current level while consuming less resources. That will allow them to maintain their trajectory until (I'm expecting) they get the huge infusion of extra compute resources from Space X later this year.

    If I'm right, then I expect we should see Google start pushing forward again (rather than more of this lateral stuff) by the end of the year.

    • nateb2022 16 hours ago ago

      > Judging from my own frustrating attempts to use Gemini for vibe-coding, it seems like Google is badly over-sold (i.e., under-provisioned).

      Based on my experience lately with Codex, it seems the opposite. Gemini 3.6 Flash (high) in Antigravity CLI feels a LOT faster than GPT 5.6 Luna (high) in codex. In the order of 10-50x.

    • addaon a day ago ago

      > putting AI queries into ever search response

      There's no way these are using a significant amount of compute. I'm not 100% convinced they're actually LLM-generated rather than an old-school Markov model. Both the relevance and accuracy numbers of the responses flirt with 0%. It's possible they just have a few million stashed responses and choose one at random, from what I can tell as a user.

  • simonpure a day ago ago
  • hmate9 a day ago ago

    This model is not for builders and engineers. DeepSWE score of 49% is behind gpt 5.4 and muse spark. It's clearly intended to be an efficient model for google gemini usage.

    What is interesting is how this is announced before any Gemini Pro progress. From the outside it seems as though Google cannot keep up with other frontier models.

    • HarHarVeryFunny 9 hours ago ago

      Presumably 3.6 Flash is primarily meant to serve their own needs for the Gemini chat app, voice app (which Sergey Brin says he uses a lot in the car on the way to work), and for their search "AI Assistant".

      Flash 3.6 is certainly capable of many coding tasks, but clearly a model of this size of not trying to compete at the frontier as a software development tool, and not clear why Google really need to complete there other than for PR-related AI bragging rights.

      I don't know how a frontier model like GPT 5.6 or Fable could have done better (I have no need that justifies paying for them), but yesterday I used the free Gemini chat app (i.e. Flash 3.6) to discuss and explain this poorly written recent AI paper to me, and honestly couldn't ask for much more.

      https://alignment.openai.com/measuring-reward-seeking/

    • andai a day ago ago

      Who tested it on DeepSWE?

      Edit: Oh it's in the other link

      https://blog.google/innovation-and-ai/models-and-research/ge...

    • thisisauserid a day ago ago

      Everyone wants to announce as late as possible (i.e., last) to chart the highest. Google is in a position financially to take a hit for these last few months.

  • m_ke a day ago ago

    I assumed google would lean into the efficiency stuff more and try to eat the easy 80% of workloads, winning market share on volume instead of frontier if they were not able to produce frontier level models.

    They're very well equipped to be the volume discount store of inference.

  • dvduval a day ago ago

    I find their models are pretty good for answering every day questions, and with their large user base, I’m sure they’re getting lots of usage.

    • handzhiev a day ago ago

      They also provide the most usage for free and in the cheap paid plans.

      • polski-g a day ago ago

        They're cannibalizing their own search ad revenue. Very strange decision.

        • HarHarVeryFunny a day ago ago

          IMO they are being pretty smart about this - the conventional wisdom is to cannibalize your own products before someone else does, and LLMs are obviously a major threat to search, with ChatGPT probably the biggest threat.

          Google's "AI Overview" search results, which started out awful, are now much improved - can be actually useful - as long as you are not asking specialist questions, and the Gemini chat and voice apps are also great for everyday use, although not sure if they are yet monetizing this (volume probably a lot less than search I'd guess).

          The Apple Siri-Gemini deal is also a significant way for Google to not get sidelined.

        • CWuestefeld a day ago ago

          In the bigger picture, people were starting to catch on to the fact that Google's search product was becoming increasingly useless.

          Around here we'd come to that conclusion at least a couple of years ago, due to abusive SEO and so forth. And that understanding was becoming even more widespread.

          I don't know if Google's got the later parts of the game figured out yet, but I have to think that they'd realized that Search was dying. While there's still some value there, may as well use it as a hook to pull people into what they expect to be the next era.

        • handzhiev a day ago ago

          They have no choice - if they don't offer AI assisted search, someone else will eat their lunch.

  • __natty__ a day ago ago
  • jgbuddy a day ago ago

    It is both less intelligent and more expensive than GLM-5.2, while being closed weight.

    • andai a day ago ago

      Their niche is that you get the quality of a Chinese model for the price of an American model.

      (Quoting myself from 2 days ago.)

  • simonpure a day ago ago

    It's available in Antigravity as well -

      Gemini 3.6 Flash (High)
      Gemini 3.6 Flash (Medium)
      Gemini 3.6 Flash (Low)
      Gemini 3.5 Flash (Medium)
      Gemini 3.5 Flash (High)
      Gemini 3.5 Flash (Low)
      Gemini 3.1 Pro (Low)
      Gemini 3.1 Pro (High)
      Claude Sonnet 4.6 (Thinking)
      Claude Opus 4.6 (Thinking)
      GPT-OSS 120B (Medium)
    • guessmyname a day ago ago

      Antigravity is such a dumb name. An-ti-gra-vi-ty (5) is long, Ge-mi-ni (3) was better.

      Aren’t Marketing/PR people at FAANG supposed to be among the best in the industry?

    • eevmanu a day ago ago

      Not yet in Antigravity CLI 1.1.5 for Antigravity Business

  • hankbond a day ago ago

    I would be fine lauding Gemini models if the only benefit of them was superior understanding of intent (read between the lines). I don't need it to code because other models are tuned for that explicitly, but I would like a model that is tuned to produce less mechanical output.

  • prometheus1992 a day ago ago

    I think there is another angle here - maybe, just maybe google doesn't want to release a pro/stronger model sooner because of two reasons- 1) they're afraid they'll have to get their hands dirty in a certain war?? 2) what's the benefit of being on top of this chart?

    so I think they're just focussed on releasing whatever helps their bottomline (improving search?). its not like anthropic and openai are making a lot of money being on the top of charts! this might just be my crazy pills talking though.

  • handzhiev a day ago ago

    Flash 3.5 does OK for various tasks. It's not super smart but is a workhorse and if 3.6 one is better than it, that's a positive for me.

  • punkpeye a day ago ago

    I just saw it in agy and started using it without asking any questions. I did not notice any major difference.

  • polbatllo a day ago ago
  • xnx a day ago ago

    Most conversation is taking place in this thread: https://news.ycombinator.com/item?id=48993414

  • lanthissa a day ago ago

    at least openai had the guts to call code red and improve.

    releasing a model worse than luna is pretty bad. Its clear that internally they did not decide coding was a thing until relatively recently.

  • wronglebowski a day ago ago

    It’s frankly embarrassing at this point. I’ve got free access through buying a Pixel phone and it’s not even worth using as it’s a waste of my time. Here’s my experience so far using it for basic sysadmin Linux type stuff.

    Gemini 3.1 Pro just feels a generation behind, from when models would miss easy things and make bad assumptions. Its not actively detrimental in bad way but the opportunity cost vs using something like Opus to be productive is large.

    Gemini 3.5 Flash is the most annoying model I have ever used. It loves to respond in ALL CAPS like “LOOK AT THAT” for no apparent reason. I realize it’s a flash model but I will give it a basic list of tasks and the output will simply vomit “Now I will X” “Now I will Y” “Now I will Z” over and over again filling my screen with garbage. It’s also no smarter than 3.1 Pro and consumes just as many tokens as 3.1 Pro, it’s really pointless without a newer Pro model in place.

    • glimshe a day ago ago

      I respectfully disagree, at least for most tasks.

      3.1 Pro: while it's coding performance is mediocre, a lot of coding work requires minimum actual thinking. I use it often for light refactoring, boilerplate generation, testcase skeleton generation, code review, language questions ("is there a better way to write this code block?"). I don't have a corporation behind me so costs matter. Considering that I'm a Pro subscriber, it's quite cost effective.

      3.5 Flash: excellent model for general (re)search. I use it for everyday tasks with Thinking instead of Flash-Lite. It's a much better version of Google Search for general queries like gaming tips, cooking, day-to-day first aid, tax and investment questions, etc.

      Google is clearly aiming for cost-benefit here and considering that it gets bundled with YouTube Plus and Google One at $20/month, it's a killer deal IMHO.

      PS: I don't work for Google and don't even like Google very much. But this is a good product.

      • Tostino a day ago ago

        It was a decent deal ~6-8 months ago. I had been using 3.1 pro almost since release, but it really is feeling old. After using other models more in the past two months though...I really can't go back to 3.1 pro, as I just have to explain my reasoning so damn much to get it on the right path, where as opus or fable just "get it" from the context of the project much better.

        Sonnet is roughly the same level as 3.1 pro for me.

    • cyanydeez a day ago ago

      This would be more useful if you could compare to a local model like Qwen3.6 27B or 35B.

  • manideep1428 a day ago ago

    where is gemini-3.5-pro ??

    • lousken a day ago ago

      my question exactly - it was announced months ago on IO

      • verdverm a day ago ago

        yeah, I was surprised when Sundar said they will release it in a month, since training/evals is not so predictable

    • antiloper a day ago ago

      You're not supposed to ask this question.

      • staticman2 a day ago ago

        It may be a dud. The blog post assures us they are now pretraining Gemini 4.

    • TiredOfLife a day ago ago

      In june. Year unspecified

    • verdverm a day ago ago

      have we defined what ai limbo is yet?

      rumor has it that pro is not so pro in internal evaluations and that there are internal stakeholders holding it up

  • alunchbox a day ago ago

    I've heard almost nothing about Gemini in my circles, are they still in the race to build AGI?

    • kingnothing a day ago ago

      Of the models I've tried lately, I get more value out of Gemini for personal stuff like researching products or learning how to do things than the others. It seems more factually correct in domains I know about and hallucinates minor stuff less than the others.

      • dzhiurgis a day ago ago

        I flop between 3.1 and 3.5 when I need it to be flawless. Fast and relatively affordable. Up to 6x cheaper than competitors.

    • HarHarVeryFunny a day ago ago

      DeepMind is absolutely still in the race to build AGI, but they see AGI as more than just a big LLM.

    • spwa4 a day ago ago

      Google has not had a single SOTA model since about a year now.

      • staticman2 a day ago ago

        You are exaggerating. 3.1 Pro came out in February of this year.

        • spwa4 a day ago ago

          ... and it was not SOTA at the time of release. Gemini 3.1 Pro was previewed on 19 February 2026. It just barely beat GPT-5.3 and was roughly equal with Sonnet 4.6. Better according to some, worse according to others. And that lasted about 10 days (until GPT-5.4 came out which was also not a huge jump). And these are large averages. On coding or terminal Gemini 3.1 Pro was not close to Opus 4.6. Also it only matches the GLM 5 open model, more or less.

          Last time Google had a "everybody agrees" SOTA model was Gemini 3 in November 2025, it beat GPT-5.1, it really was better and held it for a month.

          Currently Google's best model doesn't match open models, in any category (performance, price or speed), in fact the last 4 open weight champions all beat Google's best model.

          Here's a plot of the evolution: https://www.reddit.com/r/LocalLLaMA/comments/1v20g29/kimik3_...

          • HarHarVeryFunny 8 hours ago ago

            > Last time Google had a "everybody agrees" SOTA model was Gemini 3 in November 2025

            I have no opinion on whether this is true, but "It's been 6 months since Google had the best model in the world" seems a rather weak criticism!

            It's anyways been clear for a long time that people are finding value at all sorts of different model sizes and price points, and that Pareto frontier and cost-to-complete task are more important than who benchamaxxed who.

            If a model as strong as Gemini 3.6 Flash(!) had been released a year ago, then everyone would be falling over themselves calling it AGI - it is extremely capable, and free usage in the chat app is essentially unlimited.

            • spwa4 an hour ago ago

              Really? Weak criticism given that in that time Google was beaten by OPEN models? China is no longer behind Google in AI, Google is now catching up to China and open models. Every linux shop can give their customers better AI performance than Google can, at a cheaper cost than Google's flagship model (either $100k up front and marginal cost only, or about half the cost of Gemini 3.6 Flash).

              At least OpenAI and Anthropic can still tell their customers their best model is better than some Finnish student can setup in the customers' basement, even if the difference is pretty small now. But they are better. Google is not.

              Surely that's a major change in the situation. "America" is losing the advantage (although even the Chinese models are losing, see next paragraph) but Google is a big step behind 6 other players and 3 open models, so let's just say, Google is effectively behind everyone else that matters. Worse: one of the models that beats Google's best model was trained on Chinese ASICs (that aren't even 2nm. And very few I might add, couple thousand. Google's next model ... has a LOT to prove)

              Even Chinese models are losing their advantage. By which I mean that if you check how far Qwen 3.6 27B is behind trillion-parameter models, it's less than a year, down from easily 3 years. If that continues to go down, even Chinese trillion parameter models will become hard to justify (and Qwen 3.8 is being cooked up as we speak, I mean we don't even know if a new 27B is in the works or not, but everyone's very excited)

              • HarHarVeryFunny 20 minutes ago ago

                Has Google even being trying to keep up with frontier sized models, and for that matter why should they?! Even if they want to, what's the rush? Someone from Google just tweeted today that they just recently STARTED their "Gemini 4" (whatever that may be) pre-training run... In the meantime they just today announced financial results with cloud revenue up 82% YOY ... they seem to be doing just fine.

          • staticman2 a day ago ago

            I don't agree with you on what the term SOTA means. Gemini Pro 3.1 was the only frontier model that had native video input at launch. It was certainly SOTA at some tasks.

            • spwa4 16 hours ago ago

              That's my point. It wasn't a clear win. It did some things no other model did. And it sucked at coding and terminal (as in nowhere near SOTA). It wasn't a clear win.

          • a day ago ago
            [deleted]
  • uejfiweun a day ago ago

    I know that Flash is obviously no Opus or Fable. That being said, when I want a very quick response, my go-to is Gemini Flash. It's just really damn fast and generally tends to be accurate enough.

    • HarHarVeryFunny a day ago ago

      Yeah, Gemini Flash is my daily go to. It's certainly quite capable - yesterday it talked me though fixing my Ubuntu desktop (rebuilding NVIDIA driver) after an update caused it to boot into low resolution mode.

  • heyliot a day ago ago

    Why isn't there an official annoucement? And is pro dead haha?

  • kilroy123 a day ago ago

    Is this supposed to be a replacement for 2.5 flash?

  • a day ago ago
    [deleted]
  • m3kw9 a day ago ago

    I don't like how it compared itself to 5.6 Luna instead of Sol, and still losing on some metrics. You will not see me using this over even 5.6 Sol-light

    • HarHarVeryFunny a day ago ago

      Claude's free usage is so limited that it's useless. Gemini is extremely generous - never hits a limit for me despite being used all day.

      A couple of days ago I was comparing Claude vs Gemini asking both what the size of YACC generated parsing tables for ANSI C was. Claude was very gung-ho and rushed off to download a C grammar, install bison, and find out for itself (not really necessary), and then hit daily free usage limit half way through without even getting to an answer (despite this being the only thing I'd asked it that day).

  • zzleeper a day ago ago

    [flagged]