Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

(trufflesecurity.com)

34 points | by 882542F3884314B 12 hours ago ago

13 comments

  • croemer 9 hours ago ago

    In principle interesting, but I can't stand the Claude writing.

    > Keys with real blast radius

    > Here is what they unlock.

    > This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.

    > We cloned the public dataset hub end to end: every repository, every branch, every large-file object

    > The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.

    The whole post looks like a Claude artifact with random little cards.

    It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.

    • fractorial 9 hours ago ago

      I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.

      • xena 8 hours ago ago

        That takes a level of taste and craft that many in this industry do not possess.

        • manquer 6 hours ago ago

          It takes a level of taste and craft to write good content.

          It only bit of effort to write in your own voice , anyone with high school level of writing skills should be able to muster if they were inclined to do so .

    • gumby 8 hours ago ago

      Ever read a book by Yuval Noah Harari? He seems to have been writing like this since before LLMs. Perhaps that’s where they got it from.

      • jgalt212 7 hours ago ago

        reply with ASD-STE100 -> Yuval

    • huflungdung 7 hours ago ago

      [dead]

  • ks2048 8 hours ago ago

    I think "7.6 PB" is more informative than "4.4x Empire State Building heights worth of DVDs", but that's just me.

  • lorreyfum 9 hours ago ago

    Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.

    • croemer 9 hours ago ago

      I guess Huggingface datasets are easy to crawl - those datasets are hosted to be crawled. In contrast to the web at large which will be behind Cloudflare etc

  • Izmaki 8 hours ago ago

    This is fine. All of this is fine. :’)

  • 9 hours ago ago
    [deleted]
  • Ozzie-D 5 hours ago ago

    [flagged]