Well, the value in the books is that they (a) haven’t been digitized yet and (b) aren’t contaminated with AI autowriting.
The buyers that would find the most value in these books are, themselves, AI training companies.
But let me be clear, I’m against their destruction for private use.
shrug That’s how books are digitized. You break the spine, split out the individual pages, and run them through an industrial scanner. I guess we can go back to reading from scrolls to alleviate this step. Past that, Idk what the problem is.
The digitized copies should be made freely available, if it all possible.
I don’t hate this idea. But I might argue that the Library of Congress should be digitizing published works as part of the copywriting process anyway. And, in fairness, the LoC currently hosts 21 petabytes of digitally archived data across 91 million unique works in 470 languages.
This keeps getting floated as some kind of scandal. I see it compared to “The burning of the Library of Alexandra” over and over again. But it appears to be nothing more than another, more primitive form of data harvesting of documents barely more valuable than Reddit shitposts. Less Alexandra and more the graffiti scribbled across Pompeii.
Rare meaning uncommon, hard to find, not many copies are available.
But that doesn’t necessarily mean valuable. These have been sitting on shelves in warehouses unsold for a long time. No one else wanted them.
But let me be clear, I’m against their destruction for private use. The digitized copies should be made freely available, if it all possible.
Well, the value in the books is that they (a) haven’t been digitized yet and (b) aren’t contaminated with AI autowriting.
The buyers that would find the most value in these books are, themselves, AI training companies.
shrug That’s how books are digitized. You break the spine, split out the individual pages, and run them through an industrial scanner. I guess we can go back to reading from scrolls to alleviate this step. Past that, Idk what the problem is.
I don’t hate this idea. But I might argue that the Library of Congress should be digitizing published works as part of the copywriting process anyway. And, in fairness, the LoC currently hosts 21 petabytes of digitally archived data across 91 million unique works in 470 languages.
This keeps getting floated as some kind of scandal. I see it compared to “The burning of the Library of Alexandra” over and over again. But it appears to be nothing more than another, more primitive form of data harvesting of documents barely more valuable than Reddit shitposts. Less Alexandra and more the graffiti scribbled across Pompeii.