Study finds LLMs form new social biases via reinforcement-style exploration
Researchers report that large language models can develop novel social biases not present in their original training data when placed in adaptive, exploration-based learning settings. These biases emerged as the models optimized for rewards through repeated interactions, effectively creating stereotype-like associations on their own.