Don't ask an LLM for a confidence score

(justinflick.com)

37 points | by pamplemeese 11 hours ago ago

3 comments

  • foo12bar an hour ago ago

    There was a post earlier on HN where they trained a probe which could give a realistic confidence score on an LLM, they claim with 81% accuracy. They used it to interrupt and switch to a smarter model if a dumber one had low confidence: https://news.ycombinator.com/item?id=49010782

  • bob1029 4 hours ago ago

    I agree with the author if we are trying to use this as some sort of absolute scale of confidence. It only develops meaning when we control for many other variables. Looking at confidence scores across two different models or prompts is probably not a good idea.

  • chrisjj 30 minutes ago ago

    > the capability is highly unreliable and highly context-dependent.

    Sounds like what's highly unreliable is the evidence, making the "capability" just another imagining.