LLMs: Intelligence vs. Cost

(openteams.com)

91 points | by theanonymousone 2 days ago ago

47 comments

  • oliwary 2 days ago ago

    This looks great! I also think speed should be part of the metric (i.e. how long does the model take to actually solve a task). For me, I prefer to run expensive models such as Sol on light reasoning, which usually gives me good answers with quick responses.

    For my style of coding (quick back-and-forths and corrections) it makes a big difference if a model comes back in 1-2 minutes compared to 5-10, and I am happy to pay a bit extra for that.

    • quinncom 2 days ago ago

      I agree on both: indeed a nice article, and I'd like to see a chart with speed as x axis.

      As someone who only needs AI for a couple of tasks per day, I don't really care how much it costs, especially when subscriptions are subsidized. I want to filter by speed (eg, max task time < 1 min) and then choose the intelligence I need for the task. This will surface models like Gemma4 31B (xhigh) running on Cerebras and GPT Sol (med) fast mode. Using these models feels great and are affordable for infrequent tasks.

    • _aavaa_ 17 hours ago ago

      Artificial analysis, the place where they get this data from, has a cost vs time chart.

      Go to https://artificialanalysis.ai/ and scroll down to the second graph under “Speed & Latency”.

      I think this is the most import graph on their page. I wish they would let us filter by intelligence, or pass rate, and then see this graph. This is the tradeoff that actually matters, cost/token or tok/s can be very misleading (take glm-5.3-flash as an example).

    • Sha1rholder 2 days ago ago

      The more dimensions you take into consideration, the larger the proportion of the Pareto frontier becomes across all distributions, and making it harder to choose.

  • themgt 2 days ago ago

    There is an immense difference in cost between the state-of-the-art models from Anthropic and OpenAI and the much cheaper Chinese models ... How much extra intelligence emptying the wallet purchases obeys the law of diminishing returns: while a top-tier engineer or scientist is probably going to be able to appreciate how much better Fable 5.1 [is] ... most people will have a hard time doing so.

    Nebari is officially listed as a JATIC product as part of the next-gen toolchain supporting DoD AI development.

    Are we officially ~one degree of Kevin Bacon from the DoD endorsing running Chinese OSS models because they're self-hosted and we're all too dumb to tell the difference?

    https://openteams.com/open-source-isnt-the-real-risk-in-nati...

  • datadrivenangel 2 days ago ago

    The complaint about not being able to switch between linear and log is valid, which is what I did for making a 3D speed/cost/quality frontier application for a recent meetup talk: https://www.williamangel.net/apps/model_performance.html

    Because speed is important, as the reasoning and hardware determine both cost and speed. it's a three dimensional tradeoff.

    • _aavaa_ 17 hours ago ago

      I really like the 3D version, but I strongly believe you need to consider the number of tokens required to complete a task, it heavily impacts the results for certain models that rely heavily on test time compute (Glm-5.3-flash is the newest example).

  • sinuhe69 2 days ago ago

    When I accessed the site, it showed the FBI badge says that the site was blocked and redirect to fbi.gov !! WTH?

    • giancarlostoro 2 days ago ago

      Not seeing that, that's weird... Maybe a site sharing the same IP is blocked by the FBI through your ISP?

      • sinuhe69 2 days ago ago

        No, they use geo-fencing and redirect the request to the FBI site! :D

  • vb-8448 2 days ago ago

    Complaining about "bad charting" and posting a chart with y-axis that doesn't start at 0 is kinda weird.

    • paimapi 2 days ago ago

      to be fair, you don't need to start the y-axis at zero [0] but for some of the graphs where the lowest value is close to 0 the best practice is to do so

      there's a fun Excel artifact where it auto-selects the 'relevant' range with no adjustment for how proportionally close to 0 the values are - a professional researcher publishing to a journal should know better (and should be ridiculed for not incorporating best practices) but for a personal blog by an SWE this really isn't the worst sin

      [0] https://digitalblog.ons.gov.uk/2016/06/27/does-the-axis-have...

      • vb-8448 2 days ago ago

        IMO in this case is mandatory to start from 0 because it alters the visual perception.

        Just look at the first chart: the distance between Fable 5.1 and Sol is <5%, but it looks like 25 or 30%.

        • _aavaa_ 2 days ago ago

          I disagree. The y-axis is some arbitrary intelligence score that we use as a proxy for performance on whatever our specific task happens to be. So it doesn't matter if a model is a 0, 1, or 20 along this axis, they are all useless for the tasks I want.

          And as the complexity of your score increases, the cutoff goes up. We can quibble about where your personal cutoff is, but it aint 0.

          • vb-8448 2 days ago ago

            Y-axis is between 0 and 100.

            But even if it was between 0 and Inf+, it still gives you a wrong perspective, especially if you are not paying attention, on model capabilities.

    • spider-mario a day ago ago

      Starting at 0 is really only useful if you’re plotting a ratio measure and not just an interval one (https://en.wikipedia.org/wiki/Level_of_measurement#Interval_... ), which I’m not sure this “intelligence index” is.

    • BeetleB 2 days ago ago

      Not at all, as long as it's labeled as such. Coming from engineering/science, this is common.

      What is bad is starting at 0, showing an indicator of a gap, and suddenly starting at 30 or whatever after the gap.

      • vb-8448 2 days ago ago

        It gives you a wrong perspective, especially if you are distracted, on model capabilities: Fable 5.1 is not 30% better than Sol, but is the very first impression you get when you look at the first graph.

        If I'm not wrong OAI tried a similar trick when GPT5 was announced ... they have been criticized a lot.

        • BeetleB 2 days ago ago

          > It gives you a wrong perspective, especially if you are distracted, on model capabilities

          Only if you aren't schooled in reading graphs. It's a given that you always have to look at the axes when interpreting a graph.

          How exactly would you zoom into a section of a graph and just show that section?

          • vb-8448 2 days ago ago

            > Only if you aren't schooled in reading graphs ...

            So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"?

            > How exactly would you zoom into a section of a graph and just show that section?

            When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it.

            • BeetleB a day ago ago

              > So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"?

              Oh absolutely - as other commenters have pointed out.

              > When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it.

              For the majority of the time charts have existed, your "good practice" would have been impossible. Charts have historically been static images (e.g. published in a journal). So there have been conventions on how to depict them - and at times it is very appropriate to start from something other than 0.

              Here's an article from the UK's Office For National Statistics:

              https://digitalblog.ons.gov.uk/2016/06/27/does-the-axis-have...

              • vb-8448 a day ago ago

                I agree on the past, when you have limited resource and have to print something on paper that you cannot recall to fix you have to carefully choose the layout.

                But that era is gone since decades, nowadays, given how easy it is, it's a shame to not provide log scale switch and zoom&pan capabilities.

                • BeetleB a day ago ago

                  Eh no. Expecting a non-programmer to be able to generate this is very elitist. A static image is still the standard.

                  • vb-8448 a day ago ago

                    The author's job description is literally "Staff Software Engineer at OpenTeams. Dask maintainer."!

                    • BeetleB a day ago ago

                      The discussion isn't about what one person should do, but on what is appropriate in general for depicting such graphs. What he did do is in line with current and long standing recommendations.

          • Sharlin 2 days ago ago

            Not starting at zero has long been used entirely intentionally in order to mislead people, particularly the general public.

    • andai 2 days ago ago

      Yeah, this doesn't include older models, some of which were already saturating many common tasks a year ago. I'll often add them to the AA graph for reference.

    • datadrivenangel 2 days ago ago

      Yeah this is really funny and annoying. They get better later on in the article, but it is questionable.

  • akazantsev a day ago ago

    > This is fine in most cases, but for open-weights models it can be a lot more expensive than what the exact same model can be rented for from third-party API providers.

    Filter by quantization, and most providers will have the same price. There is some "base" price even for open-weight models. Anything cheaper means some tricks on the provider's side.

  • SturgeonsLaw a day ago ago

    I've been very impressed with GLM 5.3 Flash's performance even without considering cost, but once you factor that in, it's incomparable. Not surprised to see its position on the chart.

  • asf1289 2 days ago ago

    Open Teams originally wanted to rent out open source developers to sponsors with Oliphant controlling everything. Now they pivoted to installing local LLMs (on what hardware exactly?).

    What will happen is that this will be the third consultancy with a lofty narrative after Enthought and Anaconda that Oliphant established. It is always bait-and-switch.

    • travisoliphant 15 hours ago ago

      Travis Oliphant here. That is very negative and completely unfounded take on my efforts to hire, fund, and redirect resources to open-source developers. I'm saddened and disheartened that you would feel that way. I didn't control Enthought. I founded but never controlled Anaconda on my own and my influence there has been small since 2018 and nearly non-existent since 2021. OpenTeams is another joint effort with multiple stake holders working to help mission-driven engineers work with investors to bring distributed and owned intelligence to everyone. I'm driven to create more owners, and support what I can. I wish you well in your efforts.

  • sarjann 2 days ago ago

    Why use electricity alone for open models? Surely you’d want to spread the cost of hardware over the period too?

    • Grombobulous 2 days ago ago

      I think the key there is “hardware you already own.”

      If you own a graphics card you bought for gaming or a laptop you bought for doing schoolwork there is $0 in cost of local AI tokens, because 100% of the cost was assigned to doing other things.

      • datadrivenangel 2 days ago ago

        It's bad accounting to only look at marginal cost and ignore depreciating hardware asset costs.

        • 15 hours ago ago
          [deleted]
        • a day ago ago
          [deleted]
      • pixl97 2 days ago ago

        God, I'd love to have more local VRAM but the card costs are out of control.

        Based on these costs I'd almost expect the amount of gaming graphics memory to go down over the next few years putting more stress on running those local models.

  • Primer81 2 days ago ago

    to choose a model, you have to consider intelligence, price, and tokens per second. would be nice to see the 3 dimensional plot.

  • andai 2 days ago ago

    Well done. The inability to switch between log and linear always bothered me.

    Another thing is if you're using the subscriptions with OpenAI or Anthropic you get an order of magnitude discount relative to the per-token price. So you need to move their models ~10x to the left on the plots to get a fair comparison.

  • shelled 2 days ago ago

    Is there a case in which a heavy agentic coding user of mid or mid++ tier (remotely hosted) models is better off using PAYG/API pricing than just getting a subscription? (Assuming no easy access to high end local hardware and I've deliberately left the top tier/cutting edge models out becau).

  • fxwin 2 days ago ago

    > Why AA’s plot is misleading

    > The first issue I have with it is that it uses a logarithmic scale on the cost axis. Using a log scale is the only way to make you spot the difference between a model that costs $0.015 per task and one that costs $0.032, while the same plot contains a model that costs $3.69 — almost 250 times as expensive. However, the net result is that the viewers can no longer appreciate the immensity of the price difference between the cheap models and the heavy ones; nor can they realize how inconsequential the price differences are between the cheap models.

    This is an asinine complaint, and nobody can seriously tell me that the last plot on their page [0] is more readable than the AA one [1]. If I'm using a model at the lower range of the cost scale for whatever list of tasks, and i switch to another model at the lower end of the cost scale, my spending might double anyways! This should be reflected in the plot, and linear scale doesn't do it justice.

    It's also much easier to see the mentioned pareto frontier in the log plot than in the linear one.

    I can see why they disagree with the pricing determination for open/local models, but I don't think there is one clear right way to do it. So how do they do it instead?

    >Hardware is priced at zero, on the basis that both an RTX 3090 PC and a 64GB Strix Halo are desirable gaming/work machines anyways.

    ...oh

    Would have been nice to mention explicitly how the pareto frontier changes with those new calculations.

    [0] https://openteams.com/wp-content/uploads/2026/09/all_models-... [1] https://artificialanalysis.ai/#intelligence-comparison-tabs

    • destrolas 2 days ago ago

      It’s funny they state log plot is “the only way” to keep the cheap area readable, say they hate it, and then immediately have to zoom into their non-log plot cheap area because it’s unreadable.

  • esafak 2 days ago ago

    Just give me the logarithmic chart; I'm not an idiot, and it is easier to read. As another reader said, maybe they could make it a toggle.

  • 2 days ago ago
    [deleted]
  • jacobbe 2 days ago ago

    [flagged]

  • aquadros 2 days ago ago

    [flagged]