Ask HN: Have LLMs Plateaued?

4 points | by leandrobon 11 hours ago ago

11 comments

  • moizrocky1 an hour ago ago

    Been using Kimi, and yeah... the American ones seems like heading nowhere new.

  • markmatsushima 6 hours ago ago

    I think the giant leap in AI technology was the Transformer. ChatGPT is based on it, with an incredible chat user interface. In terms of technology, there has been no similarly significant innovation since then. But user interfaces and training data keep improving, which continues to increase our productivity.

  • wmf 10 hours ago ago

    Obviously not. And if you look at the things LLMs do poorly there's clearly plenty of room for improvement.

    • kazinator 9 hours ago ago

      LLM shill argument in a nutshell: "while there is room for improvement in catalytic converters, they are improving practically by the month! Within a decade, internal combustion engines will put out air that is so breathable, it could be used for ventilating a maternity ward."

  • kazinator 9 hours ago ago

    I suspect you could plug GPT4 from 2023 into the integrations and workflows and it would be about the same.

    The "models are getting better with every breath you take" is just an unproven claim that the AI companies want everyone to believe and want the evangelists to use in every argument.

    The AI companies have to put out new models regularly, not because of any improvement but because the failure to do so will signal stagnation. That's how it works in the tech industry. It's not like the wine or cheese industry where you just follow your hundreds-of-years-old recipe exactly and people buy the product because of that.

  • aborsy an hour ago ago

    I hope so!

  • tripleee 8 hours ago ago

    I think the big change was agents. The LLM improvements after that point have been relatively minor in impact.

  • kimjune01 9 hours ago ago

    nowhere near plateau, but right at the inflection point of diminishing returns imo

  • kypro 3 hours ago ago

    AI is beginning to do PHD-level math.

    If we're nearing the point where you can spin up 1,000 agents to look for ways improve existing models, then automatically run experiments to validate those ideas, we're more or less at RSI.

    As always with AI progress, compute will bottleneck this early on, but a few efficiency improvements could dramatically increase this pace of progress.

    I suspect we are at most 24 months from FOOM, but I suspect within about 6-12 months most frontier AI labs will be claiming the majority of their AI research will be AI-driven.

  • cyanregiment 8 hours ago ago

    Definitely, yes, in the way the calculator did.

    This is the technology - this is what it is. We've all seen it at this point. We know what it can and can't do.

    Sure, there's Mistral 7b running locally on a Macbook vs some large frontier model with MoE, context caching and a nice UI - but it's not all that different.

    Don't we know what LLMs are at this point?

    But it's not like the mystery of AI is fully solved and all opportunity has ended. Now we have to see what can be done with it - that part I think is still mostly unexplored.

    People are trying things in the software space, in music, video, legal, and there are quite a few "AI for AI" companies (AI analytics, infra solutions) - so we'll get some innovation there.

    "Harnessing" - People will figure out better ways with model routing, multi-modality, caching, combining things we already know.

    I'm slowly seeing the philosophy of mind creep in and make contact with software engineering. Somewhere between "AGI" and Chalmers' Hard Problem of Consciousness we'll get some new innovation around the concept of thinking itself and what it means to be an intelligent being.

    Autonomous driving is paving the way (hehe) for autonomous humanoid robots. A concept I originally scoffed at, but now consider to be a very likely thing to happen sooner than we might realize. This is going to change so much, it alone is enough reason to say "AI hasn't peaked" (and these will involve interdisciplinary models like vision, navigation, and LLMs).

    It's like asking if "programming has peaked" in 1999 and it would kinda be a "Yes". And then a couple years later ColdFusion would hit the shelves in a giant box like a Deluxe Edition RPG - if you need a laugh: https://www.reddit.com/r/webdev/comments/1o26h3l/i_have_deve...

  • AnimalMuppet 10 hours ago ago

    I'm not answering your actual question, but... even if they have totally plateaued technically, there's still some more improvement left to be had from people learning how best to use them (and when not to).