Risk of Gastric Cancer: Can LLMs help our patients?
Keywords:
ChatGPT, gastric cancer, large language models, unsupervised learning, supervised learning, machine learning, tokensAbstract
Purpose: Our study investigates five large language models and their ability to communicate with Asian American patients about gastric cancer.
Methods: We posed a series of questions to 5 chatbots (ChatGPT-3.5, ChatGPT-4o, Gemini, Claude, Coral) regarding gastric cancer and its incidence among Asian American subpopulations in six different languages. Using a two-way Analysis of Variance (ANOVA), we analyzed the differences in the AI self-rated score, the human evaluator score, and the manual scores per model and language to see if there were any significant differences in questions to ChatGPT.
Results: Claude, Gemini, and GPT4o showed a significant difference in performance when compared with Coral across all of the languages tested (p adjusted = 0.008, 0.038, 0.008, respectively). Amongst these chatbots, we found that Claude and GPT4o outperformed GPT3.5 (p adjusted = 0.038, 0.007, respectively). T-tests across all six languages revealed significant differences between Chinese versus Punjabi and English versus Korean, Punjabi, and Vietnamese (p adjusted = 0.042, 0.027, 0.025, 0.025, respectively).
Conclusion: Further tunability and training with larger, more diverse datasets would allow these LLMs to address information about gastric cancer health disparities within Asian American subpopulations. Moreover, LLMs may also play a crucial role in bridging education gaps within these communities, enhancing overall health literacy, and contributing to better health outcomes.
Published
How to Cite
Issue
Section
License
Copyright (c) 1969 Gloria

This work is licensed under a Creative Commons Attribution 4.0 International License.
