AI 'digital twins' are just 'funhouse mirrors' of the people they're based on

Publicly released:
International
Photo by Andres Siimon on Unsplash
Photo by Andres Siimon on Unsplash

Digital twins (artificial intelligence (AI) chatbots such as ChatGPT, DeepSeek, and Gemini, trained using extensive, specific information about individuals) don’t really perform like the people they’re based on, according to international researchers. The team compared how digital twins and their human counterparts responded to questions on 164 outcomes, and found the twins were only slightly more accurate at predicting their corresponding person’s responses when compared to a generic chatbot persona. So, no Black Mirror just yet - the team identified five key distortions in the digital twins they say make them more like 'funhouse mirrors' of the people they are meant to represent.

News release

From: AAAS

Digital twins perform more like generic LLM personas than the humans upon whom they are based

Science Advances

Digital twins are Large Language Model (LLM)-based individuals that are created using extensive, specific information about individual humans. But when Tianyi Peng and colleagues compared how digital twins and their human counterparts responded to questions related to 164 outcomes, they found the twins were only slightly more accurate in predicting their corresponding humans’ responses when compared with a generic LLM persona. What’s more, Peng et al. identified five key distortions in the digital twins that they say make them “funhouse mirrors” of the humans they are meant to represent. Researchers across the social sciences would like to use digital twins in studies because they can be surveyed repeatedly and deployed in potentially harmful simulations that would pose an ethical risk to people. But do these twins accurately reflect their human counterparts, even after being trained on more than 500 questions answered by their humans? In 19 novel, preregistered studies testing diverse outcomes (including intentions to share misinformation and perceptions of online privacy, among others) involving 1,784 unique participants, the twins’ answers correlated more strongly with answers by their humans than did the answers of more generalized demographic personas, but these correlations were modest overall. Peng et al. studied five reasons why the digital twins may have failed: the twins were not sufficiently individuated by the human data they were trained with; they tended toward stereotyping (answering as an average “woman” rather than an individual woman); they suffer from representation bias, tending to be more accurate for participants with high education and income; they contain ideological biases, such as being more optimistic than people about human behavior and technology; and they are hyper-rational compared with human participants.

Attachments

Note: Not all attachments are visible to the general public. Research URLs will go live after the embargo ends.

Research AAAS, Web page The URL will go live after the embargo lifts.
Journal/
conference:
Science Advances
Research:Paper
Organisation/s: Columbia University, USA
Funder: This research was partly funded by Columbia Business School’s AI in Business Initiative.
Media Contact/s
Contact details are only visible to registered journalists.