Bots are currently scraping the internet for LLM training data at unprecedented rates[1][2][3], driving up costs and destabilizing public-facing websites. I want to talk about how this has been particularly difficult for wikis, and has gotten much worse in the last few months.
That compressed database would have to be updated frequently, I don’t know how well it’d work.
Wikipedia has started striking deals with AI companies as a means to cover the cost of their lookups which, as much as I may not like it, is totally fair.
A wiki for some obscure indie game won’t have the same leverage.
They are.