Anthropic and OpenAI are both facing uncomfortable questions from some large AI customers over concerns about how proprietary data may be used to train AI models. Some companies are so worried that they have begun demanding assurances about how their data is handled or going so far as to place limitations on which models their employees can use, and for which tasks, The Information reports. They fear that models may be trained on their intellectual property and information.
The issue can be traced back to a June change by Anthropic. Following the change to its flagship Fable model's policies, Anthropic can now retain customer data. The company argues that it only does so to ensure that Fable isn't being misused. But some companies have raised concerns that it means sensitive business data will be caught up in the sweep.
While both OpenAI and Anthropic point out that they don't train their models on the information given to them by companies with specific enterprise contracts by default, that doesn't tell the full story. Both companies do collect metadata from the same corporate customers, and while information on exactly what that metadata contains is hard to come by, OpenAI notes that it's only used “to better understand how our services are used." Anthropic also argues that any data it collects about how customers use its products is aggregated and anonymized. And that metadata isn't used to train models.
Latest Videos From Tom's Hardware Watch full video here:
Regardless, there are still concerns over a perceived lack of clarity about what is collected. Telecoms outfit C Spire has agreements with both OpenAI and Anthropic that prevent either from using its data to train models, the report says.
However, the contracts do allow both OpenAI and Anthropic to collect C Spire technical usage data. C Spire believes that includes information about what applications AI models are connected to as well as usage data. It also worries that the AI companies may collect information about what their models get up to between generating responses.
For its part, OpenAI says that it does not use this "chain-of-thought" data to train its models. But C Spire still believes it needs a better understanding of what data is being collected, the report adds. It argues that neither AI company is being clear in its explanations.
Taking the private approach
One solution to any privacy concerns could be to use air-gapped servers, something aerospace company Northrop Grumman has already chosen to do. The Information reports that the company runs open-source AI models on its own air-gapped servers rather than trusting the likes of OpenAI and Anthropic.
Stay On the Cutting Edge: Get the Tom's Hardware Newsletter Get Tom's Hardware's best news and in-depth reviews, straight to your inbox. Contact me with news and offers from other Future brands Receive email from us on behalf of our trusted partners or sponsors
... continue reading