They are being targeted by individuals specifically just to troll them. It’s not even random corporations slurping it up as a casualty, just nasty people.
This (non-) understanding of the word no or consent should land them on some kind of watch list.
Can infrastructure providers start dedicating their knowledge of scraper infra to the public domain so we all can start to respond appropriately to these toxic networks?
Did they not glare the images, to poison datasets?
“Scraping is necessary to develop good models. It’s like building a highway – some houses must be demolished, but in the end everyone benefits,” the poster wrote in a thread titled “[AMA] I scraped all of Cara.”
Except they razed the entire countryside, including the hidden alcoves, to build this “highway” of theirs.
Oh fuck that argument to hell and back.
Plagiarism machines are not a highway.
Also, highways are and always will be extremely destructive. They kill (predomantly disadvantaged) neighborhoods, they pollute, divide and lead to “deserts”.
Called it. https://lemmy.ml/post/50384280/26832027
Sloppers love it when you do the hard work of maintaining a clean dataset for them to steal.
An ironic article to write given the Forbes AI voiceover and AI article summary…



