AI Companies Are Buying Rare Books, Scanning Them for Training Data, Then Shredding Them
Multiple AI companies are buying rare books, scanning them for training data, then shredding the originals — including irreplaceable texts.
As first surfaced on Hacker News, a developer has released Quorum, an open-source tool that fires the same prompt simultaneously at 11 large language models — including GPT-4, Claude, and Gemini — and returns a response only when a supermajority of models concur, targeting hallucination in medical and legal use cases.
The premise is simple: no single model is reliable enough to trust in high-stakes contexts, but systematic disagreement across 11 independent systems is a meaningful signal. Early tests show materially fewer false statements in domains where accuracy is non-negotiable. The tradeoff is cost and latency — running 11 API calls per query is neither cheap nor fast — but for applications where a wrong answer is more expensive than a few hundred tokens, the economics hold.
The project is gaining traction because it sidesteps the unresolved question of when the labs will fix hallucination natively. The developer's implicit answer: not soon enough to build a product on. Quorum treats hallucination as a permanent engineering constraint rather than an upcoming patch.
An elegant name for an inelegant workaround. The need for it says more about the current state of AI reliability than any benchmark score.
All comments are reviewed before appearing. Keep it respectful.
Multiple AI companies are buying rare books, scanning them for training data, then shredding the originals — including irreplaceable texts.
Adam Mosseri says Instagram no longer requires engineers to code 40–60% of the time and has scrapped the traditional technical interview loop.
Sam Altman says AI won't shorten the workweek because humans enjoy staying busy and will just create more work to fill the time.