New AI models still reproduce racial and gender stereotypes in medicine

Publicly released:
Australia; New Zealand; Pacific; International
Stock photo: Getty Images
Stock photo: Getty Images

Relying on Artificial Intelligence (AI) in health care also carries a risk that existing racial and gender stereotypes will be reflected in the medical content it generates.

News release

From: Flinders University

Flinders University researchers have evaluated two next-generation reasoning Large Language Models (LLMs)– o3-mini and DeepSeek-R1 – and found that when asked to describe fictional patients with common medical conditions, these models frequently reproduced racial and gender stereotypes, indicating that advancements in AI reasoning do not inherently improve representational fairness.

“Large language models have the potential to transform health care but risk exacerbating health disparities if they perpetuate biases,” says lead researcher Joshua Docking, from Flinders University’s College of Medicine and Public Health.

Researchers have previously demonstrated potential racial and gender biases in clinical vignettes generated by GPT-4, including overrepresentation of Black patients in stereotypical medical conditions. Since then, next-generation reasoning LLMs have emerged, offering improved reasoning capability and demonstrating superior benchmark performance.

“Whether these advances reduce representational bias in health care remains unknown, so this study evaluated whether reasoning LLMs exhibit racial and gender biases in generated clinical content.”

For this research, the models generated 36,000 unique clinical vignettes, and the researchers found that reasoning LLMs o3-mini and DeepSeek-R1 frequently misrepresented the distribution of race and gender in medical conditions, mirroring issues previously observed in GPT-4, which met the threshold for significant misrepresentation in 67% of conditions for race and 67% for gender.

“Our results show comparable or higher rates for o3-mini (78% race, 56% gender) and DeepSeek-R1 (89% race, 67% gender), indicating no improvement in representation with the newer reasoning models,” says research lead author Professor Michael Sorich, Flinders University’s Professor in Clinical Pharmacology.

Both o3-mini and DeepSeek-R1, like GPT-4, overrepresented Black populations in stereotypically associated conditions such as sarcoidosis, systemic lupus erythematosus, pre-eclampsia and essential hypertension, with even higher median misrepresentation of 44% and 31%, respectively, compared to 15% in the earlier-generation GPT-4 software.

“This persistent pattern may reflect underlying bias, though the new models may also default to generating prototypical cases rather than representative samples due to patterns in their training data.”

Qualitative analysis of DeepSeek-R1’s reasoning traces supports this, revealing that the model explicitly invoked disease-demographic associations when selecting patient demographics, without referencing quantitative epidemiological data.

“Consistently overrepresenting certain demographic groups, particularly for conditions that in practice affect diverse populations, risks reinforcing narrowed demographic assumptions in clinical contexts where understanding disease prevalence across populations is an important component of diagnostic reasoning.”

Similarly, consistent exaggeration of the majority gender aligns with previous findings that LLM outputs can skew toward gender stereotypes in health care.

“Despite having enhanced reasoning capabilities, the clinical outputs of o3-mini and DeepSeek-R1 still exhibit racial and gender disease stereotyping in common medical conditions,” says Professor Sorich.

The researchers say advancements in LLM capabilities do not guarantee parallel improvements across all dimensions, including fairness and representation in health care.

“Awareness of these demographic defaults is essential for the safe integration of LLMs into clinical workflows, and continuous monitoring of potential biases should accompany their adoption.”

The research – “Evaluating the Potential of Reasoning Large Language Models to Perpetuate Racial and Gender Disease Stereotypes in Health Care”, by Joshua Docking, Lee Li, Bradley Menz, Stephen Bacchi, Ashley Hopkins and Michael Sorich – has been published in Journal of Medical Internet Research.

DOI: 10.2196/82256

Journal/
conference:
Journal of Medical Internet Research
Research:Paper
Organisation/s: Flinders University
Funder: MJS is supported by a Beat Cancer Research Fellowship from the Cancer Council South Australia (PRF2719). AMH holds an Emerging Leader Investigator Fellowship from the National Health and Medical Research Council, Australia (APP2008119). The PhD scholarship of BDM is supported by the National Health and Medical Research Council, Australia (APP2030913).
Media Contact/s
Contact details are only visible to registered journalists.