135 comments

  • zirkonit a day ago ago

    Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

    I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

    • szszrk a day ago ago

      My wife loves bulk book hauls. Those places that we frequent, work like a literal permanent discount warehouse - books in high shelves, on pallets, everywhere. Often hundreds of issues of the same one.

      But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.

      Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.

      We were taught respect for books, but not everything is worth preserving.

      • ghaff a day ago ago

        I didn't go this year but my local town library has a book sale where books go for something like $10/bag. I donate some books to them throughout the year. There are still a lot of books available on the last or second to last day of the sale. I'm sure a huge number get pulped.

      • fakedang 12 hours ago ago

        A few days back, I spent some time going through a World Book encyclopaedia set that my parents had bought for me when I was 4. It was basically an expensive investment back then into what would turn into a lifelong reading habit.

        Most of the information in it is practically worthless today unfortunately - Saddam Hussein still alive, as were Katharine Hepburn and King Hussein of Jordan, no mention of Ceres and Pluto was still a planet, and the Internet and software had barely a mention. Most articles had less information than a standard Wikipedia article.

        That being said, the way information was presented in them still outshines anything one may find on the internet today. Even the simple elements - neat diagrams and relevant images, proper sectioning and organization of the text, footnotes to other relevant articles...

        Long gone are the days when I would simply take a bowl of ice cream and a volume and just read it end to end.

    • a_shovel a day ago ago

      A lot of these books are reference titles that are outdated to the point of uselessness and/or weren't that great/interesting when they were new. Not many people have interest in or use for textbooks from the 50s.

    • laybak a day ago ago

      I have a similar thought too each time I walk past piles of discount books.

      I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are

    • sly010 a day ago ago

      Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies "saving" old books!) But that would require giving a s*t which they don't and that is the real problem imho.

      • Aurornis a day ago ago

        They legally cannot do this.

        • wareya a day ago ago

          What I would do is announce that the scans are being preserved and will donated to the library of congress or whatever other institution is legally able to hold onto stuff like this. But that would more directly tie specific companies to the practice of destroying unwanted old books, so nobody's going to do it, or even announce it.

        • 2OEH8eoCRo0 a day ago ago

          They legally cannot scan them in entirety either but they are.

          • ChickeNES a day ago ago

            Again, under Bartz v Anthropic they can scan and train on whatever they want, as long as the original is lost in the process.

        • theroadnotbacon a day ago ago

          That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?

          • Aurornis a day ago ago

            Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

            Training an LLM on copyrighted works is not illegal.

            This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

          • syrrim a day ago ago

            Google attempted to do this 15 years ago, they got sued and stopped. It turns out that tech companies occasionally do have to follow the law, you'd think people would be happier about that...

            • keeda a day ago ago

              Wait, if you’re talking about the Google Books case, Google won. Maybe they made adjustments on how they served results but they certainly did not stop.

              • ndiddy a day ago ago

                The Google Books settlement was originally going to make Google into a clearinghouse for scans of out-of-print books. The scans would have been available for individuals to purchase for a reasonable price, and libraries and institutions would have been able to subscribe to a service that would give patrons access to the full text of every book. This deal fell apart because some research libraries and authors argued this was anti-competitive, as anyone wanting to make a competing service would have to go through the same process as Google of settling a class action lawsuit. They instead wanted Congress to pass a law to free up the rights to orphaned books. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. The result is that nobody outside Google gets to see the full Google Books scans.

                • shagie 21 hours ago ago

                  Those scans are held at HathiTrust Research Center https://www.hathitrust.org/about/research-center/ https://en.wikipedia.org/wiki/HathiTrust

                  > HathiTrust Digital Library is a large-scale collaborative repository of digital content from research libraries, administered by the University of Michigan. Its holdings include content digitized via Google Books and the Internet Archive digitization initiatives, as well as content digitized locally by libraries.

                  https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._HathiTr... and https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,...

                  > Authors Guild, Inc. v. HathiTrust (2014) was a following case related to HathiTrust, a project by the libraries of the Big Ten Academic Alliance and the University of California systems that combined their digital library collections with those of Google's Book Search. The HathiTrust case differed in two primary factors which were raised by the plaintiffs: that for viewers with disabilities, they could view the scanned text through a screen reader to make it easier to read, and offering to print out the scans as replacement copies for members of the universities if they could verify their original copies were lost or damaged. Both uses were deemed also to be fair use by the Second Circuit.

                  > The subject of the copyright of orphan works – works that may still be under copyright but with no identifiable rights holder – was a significant point of debate after both this and HathiTrust. Normally, libraries have been hesitant to loan digital copies of orphaned works as libraries may be liable for copyright violations should the copyright owner step forward to claim ownership.

                  The bill on orphan works that didn't pass was https://en.wikipedia.org/wiki/Shawn_Bentley_Orphan_Works_Act...

                  https://www.hathitrust.org/the-collection/search-access/copy...

                  And there are exceptions for copyrighted works allowing them to lend them out.

                  > Protected by copyright law, but made available: Protected by copyright law but made available on a strictly limited basis in accordance with the statutory limitations including, but not limited to, Section 107 provisions for fair use, Section 108 provisions for libraries and archives, and the rights provided to registered users with disabilities. In the absence of an applicable exception, no further reproduction or distribution is permitted by any means without the permission of the copyright holder. Lawful uses of works are provided only under the following conditions ...

                  Scanning still continues. https://www.hathitrust.org/member-libraries/contribute-conte... - though it's not at the same rate as it was during google books project.

        • sly010 a day ago ago

          That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

          People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

          • Aurornis a day ago ago

            > That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

            I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.

      • nightpool 17 hours ago ago

        Both Anthropic and Internet Archive have gotten sued into oblivion by publishing companies once, there's no way they're going to do something even more flagrantly illegal

    • icantevenhold a day ago ago

      Why do they shred them at all instead of donating or selling them again?

      • andrew_lettuce a day ago ago

        My understanding is they cut off the binding for scanning. They'd have to resell it donate by the page

        • ChickeNES a day ago ago

          They have to destroy the copy either way for it to be fair use (according to Bartz v Anthropic). Cutting the bindings off is already the faster method to scan, and once you're required to pulp the original anyway, it becomes a no-brainer.

          • shagie a day ago ago

            The wording doesn't exactly say that. https://www.akingump.com/a/web/h6WFidTYyTnoNEPehPXMYu/auvY7D...

            The relevant part of the ruling starts on page 27.

                For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
                The third fair use factor favors fair use for the purchased library copies converted from print to digital.
            
            Fair use favors the destruction to avoid accidental surplus copying of the source material. It doesn't require it. If Anthropic put everything in a warehouse, they could keep it in the warehouse... but they couldn't do anything with them afterwards. They couldn't sell them as books or donate them as that would mean that it wasn't fair use and copyright infringement would have taken place once that additional copy was distributed again. At that point, fair use and economics both favor destroying the original.

            If I make a DVD copy of an old VHS tape that I own, that's fair use format shifting. I can keep the VHS tape without issue. I cannot donate it to the library or put it out in a garage sale. When I got rid of my VHS player, I threw out VHS tapes too since they were of no use and only cluttered my shelf.

      • quickthrowman a day ago ago

        A judge ruled it was OK to scan books and save the scanned copy if you shred the physical book afterwards.

        • shagie a day ago ago

          The judge ruled that format shifting was fair use. The fair use argument was enhanced because the original was destroyed. Hypothetically, one could do only the format shifting and keep the original - but the original could not be sold or donated or otherwise given away because then the format shifted copy would not be fair use.

          I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.

    • pixl97 a day ago ago

      Yea, there are a ton of people that seem they'd rather the books get lost forever than be looked at by an AI company.

    • palmotea a day ago ago

      > Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

      Yeah, now instead of old books being turned into toilet paper, we'll get turned into toilet paper. Much better.

      But Sam Altman will become richer than God, and isn't that what really matters?

      But don't worry! You'll still have access to ChatGPT until your savings run out.

      • ChickeNES a day ago ago

        Do you have an actual point about scanning the books, or are you just using this as a soapbox to rant about AI and Sam Altman?

        • palmotea a day ago ago

          Yes, you missed it. Perhaps you should read it again until you get it?

          • ChickeNES a day ago ago

            Again, do you have an actual objection to the subject of the article, or are you just mad it enriches people you don't like?

            • palmotea a day ago ago

              > Again, do you have an actual objection to the subject of the article

              I was responding to a comment, which apparently is another thing you missed. Maybe work on not doing that, instead of snarkily asking for everything to be carefully spelled out to you?

              > or are you just mad it enriches people you don't like?

              Nice strawman, pity if someone knocked it down.

              • ChickeNES a day ago ago

                Okay, so all you have are personal attacks, goodbye.

                • palmotea a day ago ago

                  > Okay, so all you have are personal attacks, goodbye.

                  Note: I made no personal attacks.

  • patall a day ago ago

    Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?

    • Legend2440 a day ago ago

      It sounds like it is about the information in the books. The titles they're looking for are all nonfiction. Not everything is on the internet, and just because it's a few years old doesn't mean it's outdated.

      Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.

      The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

      • newsy-combi a day ago ago

        The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.

      • ChickeNES a day ago ago

        > the information density of published books is a lot higher than most internet text

        I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

        • criemen a day ago ago

          They're targeting non-fiction books, so romance novels would be out.

          • ChickeNES a day ago ago

            Booksellers noticed a huge uptick in non-fiction purchases, that is not the same thing as them not targeting fiction at all.

            Edit: Actually, I have real evidence, the Bartz in "Bartz v Anthropic" is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.

          • scottyah a day ago ago

            Is that a policy change after o4 got a little out of hand?

      • patall a day ago ago

        I see. So it is much less about 150 year old fiction books, but more about 1980s science literature that was only ever printed five times. That makes a lot more sense than what the public debate seems to be about.

      • zardo a day ago ago

        Also the ability to set the training input limit in the past could be useful.

    • Aurornis a day ago ago

      > but now a few thousand books are what is needed to run a successful AI company?

      They're scanning millions of books.

      It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

    • hyperhello a day ago ago

      They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.

    • layer8 a day ago ago

      Diversity. Modern books with modern content and in modern styles are overrepresented, old ones underrepresented.

    • timcobb a day ago ago

      I'm guessing the more integration tables they consume, the better they become at integration.

      I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

    • ijk a day ago ago

      Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

      There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.

    • acuozzo a day ago ago

      > Can someone explain why the old books could really be relevant.

      Lots and lots of information is not online. You'd be surprised.

    • theroadnotbacon a day ago ago

      I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.

    • cesarvarela a day ago ago

      I think a book is like one completion of the mind behind it, so in a way, this is just distillation.

    • aprilthird2021 a day ago ago

      The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

      Books are more likely to be about a specific topic or story or time or setting and be more information dense

    • nemomarx a day ago ago

      Writing style maybe?

    • 2OEH8eoCRo0 a day ago ago

      They aren't but these companies have more money than sense.

    • mistercheph a day ago ago

      > why the old books could really be relevant > it can barely be about the information in those books (that would be very often outdated),

      LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.

      They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

    • aisenik a day ago ago

      Cognition is encoded in language, they weren't brain-damaged yet. There's better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.

    • lofaszvanitt a day ago ago

      Well, most of the old books are much easier to read than the new ones. Back then people were much more educated and had a much wider vocabulary....

    • mars-or_wars a day ago ago

      [dead]

  • RobotToaster a day ago ago

    The sad part is, imagine how positive this could be if the scans were made available to the public.

    • IrishTechie a day ago ago

      Might add insult to injury for the publishers/authors though?

      • RobotToaster a day ago ago

        Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn't in print, so nobody is making money out of it.

        • nightpool a day ago ago

          Unfortunately copyright law does not have a squatters-rights exception

          • ChickeNES a day ago ago

            I've believed for years that it really should have one, or at least a "you aren't selling this to the public at a reasonable market-rate price (or via subscription, I care about access far FAR more than ownership), you lose all rights to it" regime.

      • doublerabbit a day ago ago

        A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amount that the artist received would be a pointless pittance. Why not just let it be free to be enjoyed by all?

        OCR Scanned for training, then tossed away or burnt. Great for nature.

        • ChickeNES a day ago ago

          Why would they be tossed or burnt? Tons of discarded books are simply recycled like any other paper object (that's why people keep saying they were "pulped")

  • rjh29 a day ago ago

    Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I'm not super happy about AI companies hoovering this all up and not making the scans available.

  • shoo 2 days ago ago

    Via google translate:

    > It has been discovered that online used bookstores across Japan have been receiving a surge of large orders for books since around August of this year. Interviews with these bookstores reveal reports of "100 books sold per day" and "days where sales have increased fivefold," leading to widespread speculation within the industry that the orders are intended to collect training data for generative AI (artificial intelligence). Further investigation revealed records of over 50 tons of books being exported from Japan to the United States. Is it acceptable for books to be consumed and discarded for AI training?

    [...]

    > Nippon Television investigated using "Sayari," a tool that analyzes import and export data, and confirmed records that a group company of this major Japanese book distributor exported more than 50 tons of "JAPANESE BOOKS" to the US since last year. Assuming that all the books were heavy hardcovers (calculated at 500 grams), this would amount to the equivalent of 100,000 books.

  • offsky a day ago ago

    Business plan: Use AI to write fake old books. Make a used bookstore to sell these books back to AI for training.

  • karim79 2 days ago ago

    Buy as many books as you want but please don't burn after scanning.

    • willmarch a day ago ago

      It’s more that they remove the spine of the book to more quickly, accurately, and easily scan through the pages. The book’s structural integrity is destroyed in the scanning process which is why they have to discard them when they are done.

    • _ink_ 2 days ago ago

      Unless somebody steps in and prohibits it they will do it, because somehow that's most cost efficient.

      • karim79 2 days ago ago

        Yeah. It somehow keeps the competitors at bay, apparently, from previous discussions about AI companies who burn books.

        Can we really trust these AI companies when they're basically assimilating human art and culture and everything, only to go on and destroy the evidence?

      • notaigenerated 2 days ago ago

        I don't understand why can't they donate the books somewhere. Even if they don't want to share the scans, why not give up the books?

        • secabeen 2 days ago ago

          Largely, because books are heavy, and there's limited demand for most used books. If you've ever been at a book sale on the last day, there are always lots of leftover ones that no one wants.

          There's a general sense that librarians are preservationists. This is far from the truth. Most books get pulped within a few years of printing. Librarians are constantly discarding books that don't get checked out to make space for new books.

        • piva00 2 days ago ago

          In the name of efficiency they cut the spines, rip the pages apart for easier scanning, books are literal dead wood which would cost a lot in shipping so it's easier and cheaper to just discard or burn the remnants.

          The book ceases to be a book right at the beginning of the extractive process.

        • EA-3167 2 days ago ago

          They can, but it would cost more and take longer, so they don't.

          It's not as though they particularly care, and it's not as though bad press matters with a captured regulatory environment.

    • itmm 2 days ago ago

      Isn't the "burn after scanning" due to copyright? It seems like destructively scanning was an explicitly allowed method in Bartz v Anthropic as transformative fair use.

      • 2 days ago ago
        [deleted]
    • bethekidyouwant 2 days ago ago

      What do you want them to do with the books after they’ve digitized them?

      • karim79 2 days ago ago

        Give them away. Put them on the street for people to pick up. Donate them to anyone. I can't tell if you're being serious or if your comment is sarcasm. I'm leaning toward the latter.

        • bethekidyouwant 2 days ago ago

          You want them to put them on the street in front of the warehouse? Do they post them on free stuff Facebook group after or just wait for the garbage truck?

  • YVoyiatzis a day ago ago

    Pretty much as it happened with vinyl records twenty years ago. I remember seeing photos of this guy somewhere in Brazil standing atop heaps of vinyl records which he had amassed with HDLR intention. Now books. Some of us hold on forever.

    • mistercheph a day ago ago

      [flagged]

      • ChickeNES a day ago ago

        "the engine of human progress"? I think you mean capitalism.

        > the books being burned by these misanthropic lunatics are not available in any other medium

        prove it, name one title

        > this is not about fascination with some particular mediumn of transmission

        it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.

        • mistercheph 19 hours ago ago

          > prove it, name one title

          if they were available, the labs would have purchased them digitally or pirated them

  • dofm a day ago ago

    AI firms buying and destroying the sources of knowledge is a weird way to get us to Fahrenheit 451 but maybe the USA will get there before it switches to metric after all.

  • andrekandre 20 hours ago ago

    once these books have all been scanned and destroyed, doesn't it mean later competitors (and free-as-in-foss alternatives) wont ever be able to catch up?

    it seems really sad that these books data isnt open in the first place...

  • glimshe 2 days ago ago

    I'm strongly against destroying books to train AI. That said, I'd be extremely interested in exploring the contents of Japanese books, which probably have a lot of stuff not available outside Japan, through a LLM.

  • Haven880 2 days ago ago

    2000 years ago Qin burnt books. Barbaric. Then Hitler burnt books. Barbaric. Now AI burnt books. Is legal and most cost effective ways. When Skynet happened, hard to pity barbaric humans.

  • ivanjermakov 2 days ago ago

    So they ran out of information on the internet...

    • notaigenerated 2 days ago ago

      They ran out of free information on the internet. They can't steal it just like that anymore because the providers are now aware of the value they could provide and could sue the now trillion dollars companies for stealing the data.

      They didn't bother when it was a non-profit sponsored by Elon Musk.

    • r_lee 2 days ago ago

      or more like they can't use the internet anymore because it's just slop now

  • mistercheph a day ago ago

    Before you buy the apologia that these books are not valuable or interesting: if they weren't valuable or interesting the AI labs would not be spending billions of dollars to purchase and scan them. Yes, everyone has a personal anecdote about pallets of garbage books but I have three insights for you that you may not have because you don't read or sift through pallets of garbage books:

    1) The labs don't want garbage books, they want interesting books that are rare and unique. They want high quality training data, random permutations of language style are fine, but what you want is unseen information, unseen patterns of thinking, unseen ideas.

    2) Most pallet of books contains lots of valuable and interesting works, maybe 1-3% but sorting through them takes time, money, and energy, that's why the labs are starting to purchase by the pallet, it's because they already have a fully automated process so they can always beat any bookseller small or large on cost to find the books of interest and value in a pile.

    3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don't get destroyed or thrown away.

    Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine. Don't expect the book burning to be an isolated incident, they are coming for every other form of stored human knowledge or art, and yes, unfortunately while scanning it they will have to destroy the original copy. And attacking the past is only the beginning.

    • ChickeNES a day ago ago

      I have sifted through everything from piled up junky independent book shops, to cast offs from research libraries, to dumpsters of end of life books. Most books really are not worth the paper they are printed on.

      > Most books of value don't get destroyed or thrown away.

      Of value to who? Most used bookstores are boutiques that over-curate and will happily refuse or recycle books that they deem are inferior/irrelevant. This is a big reason why I prefer Half Price Books over most any other used bookstore, they sell most everything.

      • mistercheph 19 hours ago ago

        > Most books really are not worth the paper they are printed on.

        Yeah, but like I said in the comment you are replying to, somewhere around 1-5% of books are worth the paper they are printed on and far more, and those are exactly the books the labs are trying to scan and destroy, they are not buying pallets of books and scanning them in order to scan the june 1995 tv guide for the 800th time, they want unique, interesting, and high quality text, because that's what feeds pretraining.

  • t1234s a day ago ago

    The value in these AI companies will be more in their proprietary training data than the models.

  • 2 days ago ago
    [deleted]
  • a day ago ago
    [deleted]
  • Armecera 2 days ago ago

    [dead]

  • fangspire a day ago ago

    [dead]

  • swingandamiss a day ago ago

    [flagged]

  • panny a day ago ago

    A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.

    • manarth a day ago ago

      Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.

    • aprilthird2021 a day ago ago

      I'm the opposite of an AI doomer but this is actually scary to me. Once knowledge is hollowed out like this how can we get it back?

      • MarkusQ a day ago ago

        We could go out in the world, have experiences, cogitate upon them, learn to write well, and then do so? It ain't easy, but that's the way it used to be done.

      • shagie 21 hours ago ago

        All the books that they are scanning exist in the national libraries where they were published and have been for the past century or so.

        Lets find a random conference book on Amazon... Genetics, Radiobiology and Radiology Proceedings, Mid-Western Conference. It was published in 1959. It's a rare book in that I can only find one copy of it on Amazon.

        https://www.amazon.com/Radiobiology-Radiology-Proceedings-Mi...

        Lets go find it...

        https://search.catalog.loc.gov/instances/e3fb7a94-2b25-5388-...

        There's one copy onsite at the Library of Congress in the General Collections and another copy offsite. You can get a reading card for the Science and Business Reading Room ( https://www.loc.gov/research-centers/science-and-business/ ) and request that book and read it.

        It also happens that my alma mater has a copy of the book too. https://search.library.wisc.edu/catalog/999554514302121 - it's in the stacks in Ebling Library.

        Every book that has been published, there's a copy of it somewhere. If it was published in the UK, it is in the British Library ( https://youtu.be/ZNVuIU6UUiM ).

        The books are there.

        What's happening is that these books are getting bought by AI companies and scanned. If they weren't bought by the AI company then... the book in the... I'm not even sure. It's in Portland... if inventory doesn't sell, it may instead get thrown out. Maybe this year, maybe in ten years - it costs money to have inventory that doesn't sell.

        Libraries will do book sales of books that aren't checked out frequently. https://www.booksalefinder.com . Sometimes they're donated, sometimes it's library discards.

        What do libraries do as they cull their collections?

        https://nwls.wislib.org/what-to-do-with-discarded-books/

        > Some libraries will donate weeded materials to community resale shops, Goodwill, or other resale shops. Some will take all of your unwanted books without question, but many resale shops don’t have the capacity to accept the large number of books that libraries discard. If you do find a place to take them, you’ll still need to use staff or volunteer time to get them there.

        > You’ll often hear this idea from well-meaning community members, and there are some fun crafts that can be made from discarded books. A quick online search will turn up many ideas for ways to use your books in crafts for kids, teens, and adults, including folded book art, blackout poetry, collage or prints made on book pages, wreaths and garlands made from book pages, and more. But even if you do a LOT of craft activities at your library, it’s unlikely that you could use up enough of your discards to even notice the difference.

        For example... https://reallifeartist.wordpress.com/tag/book-carving/ or https://www.cutandfoldbookart.com/cut-and-fold-method-instru...

        These are books that used book stores (and libraries) are trying to get rid of to free up room for inventory that will move or shelf space for books that people will check out and read.

        ... but the last copy of a book can always be found in the national library where it was published... though if you really want to preserve Genetics, Radiobiology and Radiology Proceedings, Mid-Western Conference from 1959 from being slurped up by LLM training, you can buy a copy and shelf it yourself.

      • mistercheph a day ago ago

        This is unironically part of their business plan: don't just regurgitate the world's information but also destroy all other sources of information.

  • tsylba a day ago ago

    Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.

    • quickthrowman a day ago ago

      Feel free to buy books by the ton and preserve them yourself. The simple fact they’re being sold by weight implies they’re not rare or unique.

      • mistercheph a day ago ago

        The simple fact that the AI labs are spending billions of dollars to acquire and scan them implies they are rare and unique.

        • quickthrowman a day ago ago

          Please provide proof for the billions of dollars claim, thank you.

          • mistercheph 19 hours ago ago

            25 million books

            scanning+ocr: 0.05 per page * 300 pgs: 15/bk

            acquisition+disposal cost: 2.00/bk

            total acq+scan+dispose = 425M

            settlement value / bk: $3,000 [bartz v anthropic]

            settlement probability, let's say 1%: 0.01

            expected settlement / bk : $30

            expected settlement total: 750M [industry-wide this is certainly an underestimate]

            legal fees: 25% of settlement total: 187M

            total 1.36B

            + operations, engineering, storage

    • Analemma_ a day ago ago

      What do you think happened to all these used books before the AI companies showed up?

  • duchanjo a day ago ago

    Could this be for AI training data?

    • ErneX a day ago ago

      It’s on the headline if you visit the link.

  • genxy a day ago ago

    It doesn't matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.

    • pfdietz a day ago ago

      How does one "consume knowledge"?

      • wccrawford a day ago ago

        Well, in this case, I imagine they mean by shredding books after scanning them. Since that's what's happening.

        • pfdietz a day ago ago

          That's not consuming knowledge, that's consuming cellulose and ink. If it were, then printing another copy of a book would be "producing knowledge".

          BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.

          • genxy a day ago ago

            When the refcount goes to zero the knowledge is consumed. Your take is overlay pedantic without benefit.

            Both definitions of knowledge creation can used. If one creates a book and it is never read, has it been produced? Isn't knowledge also its access and how widely it is disseminated?

            • pfdietz 7 hours ago ago

              > Your take is overlay pedantic without benefit.

              The benefit is that it properly mocks a misleading framing.

              > If one creates a book and it is never read, has it been produced?

              If one destructively scans a book that will never be read, has anything been lost?

              • genxy 2 hours ago ago

                Then say what you mean, instead of having coy false confusion and this addled manic line of conversation.

                Very little was gained.