AI Companies Are Buying Rare Books, Scanning Them for Training Data, Then Shredding Them
Multiple AI companies are buying rare books, scanning them for training data, then shredding the originals — including irreplaceable texts.
The discussion thread on Reddit's r/artificial community puts it bluntly: AWS just launched Fable 5, and the benchmark numbers dominating the conversation are a distraction. Buried in the opt-in data retention feature — one many engineering teams will enable by default in the name of model improvement — is a clause that routes customer data outside AWS's own security perimeter entirely. One developer in the thread calls it "an enterprise architecture shift," not a model feature.
That single architectural fact is rewriting deployment conversations at companies in healthcare, finance, and other regulated sectors. This isn't a minor ToS footnote — it's a material change in how AWS handles customer data versus standard cloud storage, and it demands a different tier of security review before any production rollout.
The raw capability gains Fable 5 delivers are real. But for enterprises with compliance obligations, a model that ships data to a third-party training pipeline is a liability. CISOs will likely be the last people in the building to sign off on opt-in.
All comments are reviewed before appearing. Keep it respectful.
Multiple AI companies are buying rare books, scanning them for training data, then shredding the originals — including irreplaceable texts.
Adam Mosseri says Instagram no longer requires engineers to code 40–60% of the time and has scrapped the traditional technical interview loop.
Sam Altman says AI won't shorten the workweek because humans enjoy staying busy and will just create more work to fill the time.