All too perfect: bias and aspiration in persona generation with LLM
Published in Artificial Intelligence Review, 2026
Abstract
Synthetic data generated by large language models plays a central role in the training and alignment process of other AI systems. However, this process also risks inheriting the structural biases of organic corpora and embedding new biases that stem from the design choices underlying the data creation process. This paper examines the systematic biases that emerge when large language models (LLMs) are tasked with generating synthetic personas. We introduce a reproducible, minimally conditioned pipeline that produced 40,000 personas, in four different languages, using two instruction-tuned open-weight generators (Llama−3.3-70B-Instruct and Qwen2.5-72B-Instruct), followed by a battery of quantitative analyses: name match-rates, KL divergence/skew for gender, age-pyramid comparisons, profession-gender intersectionals, adjective/sentiment profiling, and Proppian role classification. Our main findings reveal that persona generations are far from neutral. Models tend to focus on middle-aged, aspirational, and overwhelmingly positive (i.e., upbeat/optimistic) characters, while non-binary identities and many real-world occupations remain underrepresented. We conclude that contemporary training and alignment regimes produce a form of narrative sanitization that both flattens representational diversity and embeds normative assumptions, and propose that persona-based evaluation can serve as a scalable diagnostic of what generative systems value and prioritize when depicting humanity.
BibTeX
@article{correa2026all,
title={All too perfect: bias and aspiration in persona generation with LLMs},
author={Corr{\^e}a, Nicholas Kluge and Mallmann, Rafaela Weber and Kacz{\'e}r, David and Mai, Florian and Ilievska, Ana and M{\"o}nig, Julia Maria},
journal={Artificial Intelligence Review},
year={2026},
publisher={Springer Netherlands Dordrecht}
}
