I Benchmarked Local LLMs on the Laptop I Have

(mamonas.dev)

20 points | by konmam a day ago ago

7 comments

  • segmondy 13 hours ago ago

    From the article, "One guide this summer was literally titled “Open Weights You Can’t Run.”"

    I ran Kimi a few days ago at 1/2token per second. I only get to play with it during the weekend when i have time, but I'm certain I'll be able to get it up to 5tk/sec when I'm done in a month or two. So yeah, we can run them all locally. Folks might say it's not run if it's that slow, but feh! If you can run the best AI model locally at 1tk/sec, why won't you?

    • konmam 5 hours ago ago

      Well considering I can just use DeepSeek for pennies, not sure what I might get out of that 1 tk/sec locally.

      Maybe the question is not can you, but rather should you?

      • segmondy 9 minutes ago ago

        The answer is you should, that's what hacker news is all about.

  • dpoloncsak 19 hours ago ago

    This kinda goes along with my ancedotal findings "Local models are cool but still not quite there for consumer-grade hardware, but getting closer and closer"

    It just feels, at the moment, there's no task you'd want to throw at this over a frontier model, and while prices are subsidized you really can't compete at home

    • konmam 5 hours ago ago

      Yeah, maybe once the price wars settle down, but at this point no real reason with what you can get basically for free.

  • macwhisperer 6 hours ago ago

    yeah you realize you can turn reasoning off on the newer the models right? also try my version of Gemma-12b https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense or try my qwen 3.5-9b

    • konmam 5 hours ago ago

      Might give it a go, would be nice if there are some gains to be made. Though even then, need to think what’s my threshold where I would go this is good enough for me to use instead of defaulting to claude/openai/deepseek.