Basically a deer with a human face. Despite probably being some sort of magical nature spirit, his interests are primarily in technology and politics and science fiction.

Spent many years on Reddit before joining the Threadiverse as well.

  • 0 Posts
  • 840 Comments
Joined 2 years ago
cake
Cake day: March 3rd, 2024

help-circle







  • I can imagine a couple of reasons.

    The most straightforward one actually is “capitalism”, but maybe if I explain it a bit you’ll accept it? :) I expect that the vast majority of the books that these companies are scanning are bought in bulk from wholesalers at the cheapest possible price, since the goal is quantity over quality. That means they’re probably buying the leftover books that the wholesalers are just one short step away from sending off to be recycled into pulp anyway. So re-binding them would mean they’re still left with nearly-worthless books that nobody else wanted to buy. It’s like those book sales libraries have from time to time, they put out the books that are scheduled for “disposal” in hopes that somebody will pick up a few before they go into the dumpster.

    The other reason depends on the physical nature of the book. There’s lots of different book binding techniques and not all of them will leave you with something that’s easy to put back together. Lots of books are made up of smaller subunits of pages called “signatures”, typically 8, 16, or 32 pages, that get stitched together into the finished product. There are coil-bound books, comb-bound books, all sorts of things like that. A lot of them probably wouldn’t leave pages that are amenable to a one-size-fits-all rebinding. So that makes things a lot more costly, and again that’s a cost that produces a book that’s probably worthless for resale anyway.

    Honestly I suspect that getting chopped and scanned is the best case scenario for most of the books in this situation. It converts them to digital form, rescuing their contents from being lost. The thing that bugs me is just how these companies squirrel away their hoards of books from the public. And the blame for that lies largely on copyright law, so I can’t blame them too much.



  • The thing the headlines keep skipping over in pursuit of outrage-clicks is that this is how you scan books in bulk. The spine is removed and the pages are fed through a high-speed scanner in sequence.

    And the books that these companies are scanning this way are not “rare” in the sense of Gutenberg bibles or first-editions of famous works. These are books that have been sitting untouched in libraries or warehouses for decades because nobody wants them and their alternative fate is simply to be pulped and recycled. AI companies aren’t interested in quality, they’re interested in quantity. It would indeed be nice if they’d release the results of doing these scans to the public, of course, but that’s where copyright raises its head to ruin everything. They’re just following the law, unfortunately.

    I’m all for getting more books scanned and uploaded to Anna’s Archive. But that’s a separate issue.