Documents recently unsealed in a copyright case show employees weighing the cost of buying books to train early ChatGPT models.
‘Sketchy AF’: What to Know About How OpenAI Staff Discussed Book-Pirating
Why This Matters
Internal OpenAI communications reveal employees debated the risks and ethics of using pirated books to train ChatGPT rather than paying for licensed content, weighing costs against convenience. This matters because it exposes the tension between AI companies' rapid development goals and copyright law, fueling ongoing litigation that could reshape how AI firms source training data. The revelations could increase legal and reputational risk for OpenAI and set precedent for the broader AI industry's data practices.
Key Takeaways
- Unsealed documents show OpenAI staff internally discussing the use of pirated books for training data, with some acknowledging it was ethically questionable.
- The disclosures come from an active copyright lawsuit, adding pressure to ongoing legal battles over AI training data sourcing.
- The case highlights broader industry concerns about how AI companies balance speed and cost against intellectual property rights and legal compliance.
Get alerts for these topics