Chatbots are making our writing more similar but losing the identity and personality behind it

Publicly released:
International
Photo by Emiliano Vittoriosi on Unsplash
Photo by Emiliano Vittoriosi on Unsplash

Large language models (LLMs) are making writing styles more similar without changing the overall meaning, according to an analysis of more than 880,000 texts across different writing types. The authors suggest that this could make it more difficult to identify important clues about aspects of a person’s identity, personality, and mental health from their language. The team analysed Reddit stories, news articles, academic papers, essays, social media posts, and political speeches over time to see how the rise of widespread LLM usage has changed our writing style. They found evidence of both an immediate shock, where the introduction of ChatGPT coincided with a sharp drop in variability, and a sustained effect, in which higher levels of LLM usage continued to reduce the complexity of language over time. They also analysed writing submitted by people who had completed a questionnaire to determine their personal traits, both before and after it had been edited by large language models, finding that they were less accurate at determining those personal traits using writing that had been edited by AI.

News release

From: Springer Nature

Artificial intelligence: Writing tools make language more similar

Large language models (LLMs) are making writing styles more similar without changing the overall meaning, according to an analysis of more than 880,000 texts across different writing types, published in Nature Human Behaviour. The findings indicate that using LLMs could make it more difficult to identify important clues about aspects of a person’s identity, personality, and mental health from their language.
People express similar ideas in different ways, and these differences can provide information about them, such as their social background. Researchers use language patterns to study individuals and societies, but LLMs could reduce this variation by favouring common patterns of language.
Zhivar Sourati and colleagues analysed more than 880,000 texts, including Reddit stories, news articles, academic papers, essays, social media posts, and political speeches. They also asked GPT-3.5, Llama 3 70B, and Gemini Pro to rewrite thosands of human-written texts. LLM rewriting was found to reduce variation in writing complexity by between 21–50%. In 87% of cases, the original and rewritten texts had meaning-similarity scores above 0.95, indicating that the LLMs largely maintained the original meaning.
After rewriting, computer models used to analyse the LLM texts were also an average of six percentage points less accurate at identifying authors’ personal characteristics from the writing. The LLMs weakened some language patterns linked to these traits, including associations between pronoun use and extraversion, friend-related words and loyalty, and future-focused words and age. However, other associations remained, including those between negative-emotion words and neuroticism, religion-related words and purity, and social words and gender.
The findings suggest that LLM-assisted writing could make language-based assessments less reliable in areas including psychology, mental healthcare, recruitment, and personalized services. Further research is needed to understand why LLMs preserve some personal language markers but weaken others.

Journal/
conference:
Nature Human Behaviour
Research:Paper
Organisation/s: University of Southern California, USA
Funder: This research was supported, in part, by the Army Research Laboratory under contract W911NF-23-2-0183, by DARPA INCAS HR001121C0165, and by Air Force Office of Scientific Research A9550-23-1-0463. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of DARPA, AFOSR or the US Government. The US Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein.
Media Contact/s
Contact details are only visible to registered journalists.