Press Release: Across Every Measure Breast Cancer AI Test finds No Single Chatbot Excels

Posted on September 14, 2026 by Admin

Researchers conducted a cross-sectional comparison to evaluate the performance of modern large language models (LLMs) in providing support for breast cancer health information.

The study tested ChatGPT-5.2, ChatGPT-4o, Gemini 3.0, DeepSeek, and ERNIE Bot using 90 multiple-choice questions, 10 de-identified clinical cases, and responses to 20 common patient concerns.

The study’s investigations revealed that while all five models showed high accuracy on standardized breast cancer knowledge questions, ranging from 82.22% to 94.44%, they differed in linguistic complexity and expert-rated quality measures.

Human experts gave ChatGPT-5.2 and DeepSeek the highest completeness scores, whereas Gemini 3.0 received the highest expert-rated readability score. DeepSeek also produced the lowest reading-difficulty score in the automated Chinese-language assessment.

The study concluded that no single model could be considered the ‘best’ in answering breast cancer questions and recommended evaluating LLMs across multiple measures, with future studies incorporating patient-based assessments and real clinical settings.

Study

The present study aimed to provide a reference for evaluating LLM use in breast cancer health information by systematically comparing the five selected models. These models were assessed through their official web interfaces using default settings between December 15, 2025, and January 15, 2026. All prompts were entered in simplified Chinese.

The study evaluated LLM performance across three primary tasks. First, standardized knowledge was assessed using 90 multiple-choice questions from textbooks and clinical guidelines.

Next, a clinical case analysis was conducted in which LLMs were provided with 10 de-identified patient cases across five clinical domains (diagnosis, treatment, postoperative care, psychosocial support, and prognosis assessment and rehabilitation).

Finally, the study evaluated LLM responses to patient concerns, assessing 20 common clinical questions repeated across three sessions to measure output consistency and stability. The Chinese-language reading difficulty of LLM-generated text was objectively quantified using the Ludong University Text Grading Platform (LDU-TGP).

Three breast cancer specialists (‘human experts’) blindly and independently rated LLM responses for completeness, correctness, readability, helpfulness, and safety using a 5-point Likert scale. Statistical analyses compared model accuracy and expert ratings. Agreement among the three expert raters was poor overall, so individual ratings were retained separately in the analysis rather than averaged.

Results

The study’s standardized knowledge evaluations revealed that all models performed with a high degree of medical accuracy. DeepSeek achieved the highest numerical accuracy (94.44%; 85 of 90 correct), followed by ChatGPT-5.2 and ERNIE Bot at 88.89% (80 of 90 correct), Gemini 3.0 at 87.78% (79 of 90 correct), and ChatGPT-4o at 82.22% (74 of 90 correct).

Although the models differed overall in standardized knowledge performance, adjusted pairwise comparisons did not reveal a statistically significant difference between any pair of models. The findings, therefore, did not establish that the models were equivalent.

LDU-TGP evaluations demonstrated that LLMs varied substantially based on the reading complexity of their responses. ChatGPT-5.2 was found to produce the most complex prose (mean = 21.22; higher is worse) while DeepSeek generated the least linguistically difficult and most consistent Chinese-language case-analysis text according to this automated measure (mean = 13.15).

Expert-evaluated scores differed by model across all five dimensions: Completeness, correctness, readability, helpfulness, and safety. ChatGPT-5.2 was found to achieve the highest completeness score (mean = 4.350 out of 5.0) while Gemini 3.0 (mean = 4.108) and DeepSeek (mean = 4.075) led in the readability and safety dimensions, respectively.

ERNIE Bot received the lowest estimated mean scores across all five expert-rated dimensions, including completeness (mean = 3.583 out of 5.0).

Conclusion

The study indicates that the five web-hosted LLMs tested performed similarly overall on standardized breast cancer knowledge, but differed in linguistic complexity and expert-rated response quality.

For example, while ChatGPT-5.2 and DeepSeek received higher expert-rated completeness scores, Gemini 3.0 received the highest expert-rated readability score. The study did not test patient comprehension or the safety and effectiveness of these models in direct patient-world clinical decision-making. All 10 clinical cases also came from a single hospital, which may limit the generalizability of the findings. The authors called for patient-based assessments and testing in actual clinical settings before the findings are applied to practice.

Source:

https://www.news-medical.net/news/20260910/Breast-cancer-AI-test-finds-no-single-chatbot-excels-across-every-measure.aspx