Unsealed filings show OpenAI staff debated using pirated books to train ChatGPT
Court documents made public in an ongoing copyright lawsuit reveal internal OpenAI communications in which employees discussed the expense and legality of acquiring books to train early ChatGPT models, with some staff flagging the sourcing as questionable. The messages suggest teams weighed pirated material as a cheaper alternative to licensing content properly.