AI Companies Are Buying Rare Books, Scanning Them for Training Data, Then Shredding Them
Multiple AI companies are buying rare books, scanning them for training data, then shredding the originals — including irreplaceable texts.
A new study reported by MIT Technology Review finds that large language models deployed for résumé screening don't simply reproduce the biases embedded in their training data — they develop new prejudices of their own. The distinction matters: biases imported from human trainers can, in principle, be identified and corrected. Biases that emerge from the model's own processing are harder to predict and harder to diagnose.
The research arrives as AI-powered hiring tools have become standard across corporate recruiting. Companies have increasingly delegated first-pass screening to these systems under the assumption that removing humans from the process removes human prejudice. This study's finding complicates that assumption directly.
The upshot is that AI hiring tools may be producing different outcomes than what humans would produce — without those outcomes being more fair. "Different" and "better" are not synonyms, and the study adds weight to the argument that companies deploying these systems in high-stakes environments carry responsibility for outcomes they may not fully understand.
Whether AI should be making first-cut decisions about people's livelihoods is an open policy question. This research makes it harder to answer that question with a confident yes.
All comments are reviewed before appearing. Keep it respectful.
Multiple AI companies are buying rare books, scanning them for training data, then shredding the originals — including irreplaceable texts.
Adam Mosseri says Instagram no longer requires engineers to code 40–60% of the time and has scrapped the traditional technical interview loop.
Sam Altman says AI won't shorten the workweek because humans enjoy staying busy and will just create more work to fill the time.