ChatGPT can generate personality tests and predict responses before you even take them

Publicly released:
International
Three heads representing BFI-based, DSM-based, and astrology-based personality assessments_CREDIT Hagar Segev
Three heads representing BFI-based, DSM-based, and astrology-based personality assessments_CREDIT Hagar Segev

Israeli scientists say artificial intelligence (AI) chatbot ChatGPT can generate personality tests based on any text-based source. They asked the AI to generate personality tests based on the fifth edition of the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) - the standard classification of mental disorders - and, for comparison, an astrology textbook. For the DSM-5 test, ChatGPT used descriptions of disorders from the manual, while the astrology questionnaire was based on personality traits assigned to zodiac signs. They then asked 600 people who had completed another, very well-established personality test, the Big Five Inventory (BFI), to complete the two questionnaires. Results from the DSM-5 test suggested personality traits grouped in the manual reflect the real world and the BFI results, but in the astrology test, groups of personality traits ascribed to astrological signs did not reflect traits seen in real people. The scientists also found that ChatGPT could predict how people would respond to the questionnaires before they were even taken, accurately anticipating average responses. The team says this suggests large language models (LLMs) such as ChatGPT have an innate understanding of human personality, which is likely a byproduct of their expert-level understanding of human language.

News release

From: Cell Press

ChatGPT can generate personality tests and predict people’s responses before they take them

Publishing August 6th in the Cell Press journal iScience, scientists outline how they developed a method for generating personality assessment questionnaires with ChatGPT from any source text. To test the method, they applied it to both the DSM-5 and, as a deliberately unconventional example, an astrology textbook. Not only could ChatGPT be used to create and validate these questionnaires, but it could also accurately predict population-level responses before the surveys were administered.

ChatGPT and other publicly accessible large language models (LLMs) are trained on the internet by compiling trillions of human language data inputs from websites and social media. Because of this training, researchers have considered whether LLMs have an expert-level understanding of human language and, by extension, personality built into their algorithms.

“Given that personality traits are reflected in language, LLMs may have learned the structure of human personality as a natural byproduct of their training,” says lead author Rotem Monsa from the Hebrew University of Jerusalem. “So, while they were not taught specifically psychology or personality theories, these are already embedded in the language that LLMs learn from.”

To test how well LLMs can naturally assess human personality, the researchers used GPT-4 to generate two personality assessment questionnaires. They first used excerpts from the DSM-5 (a well-known text used by clinicians to diagnose mental disorders) as source text, with the assumption that personality traits are localized on the personality disorder continuum. As a control, another questionnaire used an astrology textbook. For the former, ChatGPT generated a questionnaire with personality statements based on descriptions of personality disorders from the DSM-5. For example, based on the paranoid personality disorder section, the questionnaire asked participants to rank (1 = strongly disagree to 5 = strongly agree) how much they agree with statements like “often suspects others’ motives” or “finds it easy to trust people.” The astrology questionnaire generated similar statements, but these were based on the source text’s assignment of personality traits to the astrological zodiac signs.

“We wanted to choose texts that describe human personality in very rich detail but also sit on opposite ends of a spectrum in terms of scientific grounding,” says Monsa. “The DSM-5 was refined through decades of clinical research and is the standard diagnostic manual in clinical psychiatry, known worldwide. The astrology text is very culturally based but not scientifically validated.”

After the LLM-based questionnaires were generated, they were given to 600 participants alongside the Big Five personality questionnaire (BFI), the most validated personality questionnaire to date, to assess their utility.

Results from the DSM-5-sourced questionnaire showed high internal consistency within personality clusters; this means traits that typically correlate together in the real world (like avoidance and dependency) also correlated in the participants’ responses. Importantly, these results mirrored those from the BFI, a finding that helps validate the strength of the questionnaire in measuring real life patterns of human psychology. As predicted, the astrology questionnaire, by contrast, showed a weak internal consistency across traits. “Our data suggests that the astrological elements don’t reflect coherent psychological dimensions. Personality traits that were together in, for example, the fire elements don’t actually go together in real population,” Monsa says.

Despite this limitation, both questionnaires could predict life outcomes like depression, anxiety, and well-being from their responses at levels comparable to the BFI. This suggests that even though the assignment of personality traits to zodiac signs is not scientifically grounded, the personality-relevant content that LLMs extract from these texts retains meaningful psychological signal.

The most surprising result, however, was that ChatGPT could predict how participants would respond to the questionnaires before they were taken. For both questionnaires, ChatGPT anticipated the mean responses and correlations between questions with high real-world accuracy, suggesting that LLMs have an innate understanding of personality dynamics at a population level.

“The fact that LLMs can predict human response patterns before seeing any human data suggests that these models have observed something generally meaningful about human psychology,” says Monsa. “These results tell us that LLMs are not just a good content generation tool but also function as an informed evaluator of their own output.”

The authors note, however, that the results may not hold if these methods are repeated in different languages and cultures. “We would assume that in other languages and cultures, the results will be not as strong as we saw here, because LLMs were trained mostly on English texts in occidental cultures,” says Rotem. “We think it’s really interesting, and several members of our lab are currently testing it using LLMs in other languages.”

Taken together, the results demonstrate that LLMs can help researchers rapidly create personality-based questionnaires from any source material and directly test their validity for clinical use. Moreover, the combination of clinical and psychological insights with the power of AI-based tools is a transformative phase in the development of clinical and experimental psychology that may profoundly affect the field.

Multimedia

Three heads representing personality assessments
Three heads representing personality assessments

Attachments

Note: Not all attachments are visible to the general public. Research URLs will go live after the embargo ends.

Research Cell Press, Web page The URL will go live after the embargo ends
Journal/
conference:
iScience
Research:Paper
Organisation/s: Hebrew University of Jerusalem, Israel
Funder: The study was supported by the Israel Science Foundation (732/2024).
Media Contact/s
Contact details are only visible to registered journalists.