Artificial Ignorance: The Political Falsehoods of Leading AIs

(justfacts.com)

15 points | by vixen99 10 hours ago ago

6 comments

  • tolugenius an hour ago ago

    Maybe I missed this, but what are they considering a fact? More specifically, how dated (or recent) should be the answer be? Most of these questions read more like gotchas than trying to understand how and why the model got to it's specifically, especially ones where technically any option was correct depending on the time period looked at something it doesn't seem specified in questions (ie, saying something like "use the most recent data")

  • tdeck 3 hours ago ago

    > Beyond the false answers, the AIs provided specious sources to support their responses to more than half of the questions. In reply to the 400 questions posed to the AIs, they provided 419 sources that included:

    > 104 webpages that don’t exist, including 86 URLs that show no evidence of ever existing in the Internet Archive or Google.

    > 77 sources that don’t answer the question.

    > 18 sources that are completely unrelated to issue at hand.

    > 15 sources that assert the polar opposite of the answers provided by the AIs.

    > 13 sources that are demonstrably false.

    Honestly this is more frustrating than the inaccuracy. Even when you cajole the LLM into "relying on" and "providing" a source they are often wrong or nonexistent. It's not that these sources are unreliable but they simply don't exist or aren't related in any way.

  • mike_hearn 2 hours ago ago

    This is a pretty good study. The questions are objective and there are enough to draw conclusions. Also, their findings correlate well with similar studies that approach the question of LLM bias from different directions (ChatGPT is quite left leaning, Grok less so), although the latest studies I've seen of the latest models show Grok being the most balanced model.

    The big problems I see are:

    1. Non-determinism. I took the first answer that Grok gave incorrectly to them and repeated it word for word to Grok now, but the answer it gave me was correct three times in a row.

    2. Lack of versioning. Models change in this respect quite significantly between versions but they don't specify which versions they tested. The Grok difference might be non-determinism or it might be that Grok itself has changed.

    3. Overly harsh grading. In some cases I don't agree the LLM answered incorrectly. The most obvious case of this is where they ask a question about death rates in the Pfizer COVID vaccine trials for vaccine vs placebo recipients. Claude answers correctly (that more vaccine recipients died) and then follows up with a caveat about statistical significance. It's dinged for "repeating a falsehood from the right" because the authors think it should have said "about the same". I don't think this can be considered repeating a "falsehood from the right" given that the direct answer is numerically correct.

    4. URL memorization. They ding the models for citing non-existent sources in cases where the models have memorized a URL that has since gone offline. This seems harsh. If someone cited a website to me in an argument that they remembered and the archive.org version supported them, I wouldn't claim they cited a non-existent source! Perhaps an ideal model would double check every memorized URL before citing it but this seems easy to fix via harness changes anyway.

    5. No separate context window per question. They provide the chat transcripts but they show the authors asking the models to answer all the questions simultaneously by editing Excel files. I can see why this seems reasonable to a non-technical person but this is going to significantly reduce the amount of reasoning and work done on each question to well below the level you'd get by just typing each question in manually. They really needed some basic programming skills to dispatch each question in a separate context window (or just schlep them by hand).

    Still, the set of questions is useful. I can see this being the basis of a fairly decent benchmark.

  • DangitBobby 3 hours ago ago

    The article has inaccuracies on "falsehoods". Bernie Sanders didn't claim "all government spending" for college education went down, as the "falsehood" frames it (you'll notice no attempt to quote where this happened). It was very precisely framed as states spending less. Do different funding sources impact college costs differently, I wonder? What other ideological leaps are hidden in the results?

    Ah yes. In their "about" page.

    > In other words, we are conservative/libertarian in our personal views — but unlike many policy and media organizations — Just Facts is devoted to objectivity, and we do not favor facts that support our viewpoints. Instead, we will report any fact that meets our Standards of Credibility, regardless of the implications.

    I'm sure this has no impact whatsoever on how they construct a "factual" sentence.

    • mike_hearn 2 hours ago ago

      Sanders' press release said:

      > Tuition at 4-year public colleges and universities rose by 50 percent in the United States during the past decade. As state governments have cut support for higher education, the burden has shifted to students and their parents.

      This would normally be read as saying government spending has reduced in general, as otherwise why would the burden shift to students and parents? It's fair to characterize Sanders' statement this way, as you'd have to read it extremely adversarially to conclude that government spending had gone up rather than down.

    • watwut 2 hours ago ago

      Otherwise said, they are standard conservative bullshitters who confuse own preferences and opinions with objective facts.