Simon Münker, Nils Schwager, Achim Rettinger
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL)
The ability of Large Language Models (LLMs) to mimic human behavior triggered a plethora of computational social science research, assuming that empirical studies of humans can be conducted with AI agents instead. Since there have been conflicting research findings on whether and when this hypothesis holds, there is a need to better understand the differences in their experimental designs. We focus on replicating the behavior of social network users with the use of LLMs for the analysis of communication on social networks. First, we provide a formal framework for the simulation of social networks, before focusing on the sub-task of imitating user communication. We empirically test different approaches to imitate user behavior on X in English and German. Our findings suggest that social simulations should be validated by their empirical realism measured in the setting in which the simulation components were fitted. With this paper, we argue for more rigor when applying generative-agent-based modeling for social simulation.
Simon Münker, Achim Rettinger, Damian Trilling
Challenging the Myth: A Research Arc on LLMs as Human Simulacra
Proceedings of The Big Picture v2: Crafting a Research Narrative at the Annual Meeting of the Association for Computational Linguistics (ACL)
When Large Language Models (LLMs) combined with prompt-based approaches as human simulacra emerged, they promised revolutionary shortcuts. Models trained on vast internet corpora may replicate human behavior and communication through text-based alignment. The initial optimism of the NLP community positioned LLMs as universal human proxies capable of replacing participants in surveys, generating authentic social media content, and simulating diverse cultural perspectives. We systematically dismantle this "myth of universal generalization" and document a shift toward methodological rigor. Our research reveals fundamental limitations: LLMs exhibit inhuman response patterns in psychometric assessments and produce detectable synthetic content. We analyze the difference between superficial linguistic fluency and genuine human-like representation, and reframe the current paradigm from asking "can LLMs replace humans?" to "under what validated conditions might LLMs serve as useful research components in social sciences?" Our work shows how interconnected research efforts challenge foundational assumptions and establishes best practices for deploying LLMs as human simulacra.
Simon Münker, Fabio Sartori
Guardrail Vulnerabilities in Open-Source Language Models: Implications for Democratic Discourse and Marginalized Communities
Proceedings of the 59th Hawaii International Conference on System Sciences (HICSS)
The proliferation of open-source Large Language Models (LLMs) presents a complex technological phenomenon with significant societal implications. While these models democratize access to advanced Natural Language Processing (NLP) capabilities, they simultaneously amplify risks for marginalized communities who often bear the disproportionate burden of technological misuse. Our research examines systematic vulnerabilities in guardrail mechanisms across seven prominent open-source LLMs, revealing patterns of harmful content generation that threaten democratic discourse and social cohesion. Through empirical analysis using advanced NLP classification methods, we demonstrate that popular open-source models consistently generate content classified as hateful or offensive when subjected to adversarial prompting techniques. These findings directly contradict the safety assurances provided by model developers, particularly Meta AI's stated commitment that their systems should present balanced perspectives on debated policy issues rather than singular viewpoints.
Simon Münker, Nils Schwager, Kai Kugler, Michael Heseltine, Achim Rettinger
Next Reply Prediction X (NRP-X) Dataset: Linguistic Discrepancies in Naively Generated Content
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) at the Language Resources and Evaluation Conference (LREC)
The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. While LLMs offer scalability and cost-efficiency, their “naive” application, where they are prompted to generate content without explicit behavioral constraints, introduces significant linguistic discrepancies that challenge the validity of research findings. This paper addresses these limitations by introducing a novel, history-conditioned reply prediction task on authentic X (formerly Twitter) data, to create a dataset designed to evaluate the linguistic output of LLMs against human-generated content. We analyze these discrepancies using stylistic and content-based metrics, providing a quantitative framework for researchers to assess the quality and authenticity of synthetic data. Our findings highlight the need for more sophisticated prompting techniques and specialized datasets to ensure that LLM-generated content accurately reflects the complex linguistic patterns of human communication, thereby improving the validity of computational social science studies.
Nils Schwager, Simon Münker, Alistair Plum, Achim Rettinger
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
Proceedings for the 15th Workshop on Computational Approaches to Subjectivity, Sentiment Social Media Analysis (WASSA) at the European Chapter of the Association for Computational Linguistics (EACL)
When Large Language Models (LLMs) combined with prompt-based approaches as human simulacra emerged, they promised revolutionary shortcuts. Models trained on vast internet corpora may replicate human behavior and communication through text-based alignment. The initial optimism of the NLP community positioned LLMs as universal human proxies capable of replacing participants in surveys, generating authentic social media content, and simulating diverse cultural perspectives. We systematically dismantle this "myth of universal generalization" and document a shift toward methodological rigor. Our research reveals fundamental limitations: LLMs exhibit inhuman response patterns in psychometric assessments and produce detectable synthetic content. We analyze the difference between superficial linguistic fluency and genuine human-like representation, and reframe the current paradigm from asking "can LLMs replace humans?" to "under what validated conditions might LLMs serve as useful research components in social sciences?" Our work shows how interconnected research efforts challenge foundational assumptions and establishes best practices for deploying LLMs as human simulacra.
Sjoerd B. Stolwijk, Mark Boukes, Wang Ngai Yeung,Yufang Liao,Simon Münker,Anne C. Kroon, Damian Trilling
Can we use automated approaches to measure the quality of online political discussion? How to (not) measure interactivity, diversity, rationality, and incivility in online comments to the news
Communication Methods and Measures
This article explores the (in)ability of automated tools to measure the deliberative quality of online user comments along the standards set out by Habermas: interactivity, diversity, rationality, and (in)civility. Utilizing a stratified sample of manually coded comments (n = 3,862) responding to news videos on YouTube and Twitter, we examined the performance of rule-based measures (i.e. dictionaries), machine-learning classifiers (conventional and transformer-based) and measurements by generative AI (Llama 3.1, GPT-4o, GPT-4T). We present results for over 50 metrics side-by-side to judge the opportunity costs of choosing one method over another. The results revealed strong variation across different groups of models. Overall, our expectation that more modern methods (transformers and generative AI) outperform the older, simpler ones was confirmed. However, the absolute differences between these model groups strongly depended on the measured concept, and we observed strong variance in performance among models of the same group. We provide recommendations for future research that balance ease of use with the performance of automated measurements, along with important cautions to consider.