Open-access Can artificial intelligence manage malnutrition? A critical look at ChatGPT's performance in geriatric nutrition

SUMMARY

BACKGROUND:  Artificial intelligence represents a rapidly advancing innovation in healthcare with the potential to revolutionize the field of clinical nutrition.

OBJECTIVE:  The aim of this study was to evaluate ChatGPT's potential to support the clinical decision-making process regarding nutrition in older adults.

METHODS:  Twelve questions and three clinical vignettes addressing fundamental concepts of malnutrition, including general information, diagnosis, follow-up, and treatment, were created and asked to ChatGPT. Three geriatricians independently examined ChatGPT's responses. The quality of the responses was assessed using the Quality Analysis of Medical Artificial Intelligence tool.

RESULTS:  The inter-rater reliability among the authors was calculated, and an excellent intraclass correlation coefficient of 0.84 (95%CI 0.77–0.89; p<0.001) was found. The total mean Quality Analysis of Medical Artificial Intelligence score for the ChatGPT-generated responses to questions related to malnutrition was 26.60, indicating very good quality. In evaluating the clinical scenarios, the lowest scores were observed in source use. The total Quality Analysis of Medical Artificial Intelligence, accuracy, relevance, and use of sources scores for the clinical scenario involving the patient with a hip fracture were statistically significantly lower compared to other scenarios.

CONCLUSION:  Our study highlighted that ChatGPT has the potential to generate correct answers related to complex clinical scenarios about malnutrition. ChatGPT can help clinicians make more informed decisions regarding patients’ nutritional requirements and management by utilizing more up-to-date medical resources and guidelines.

KEYWORDS:
Artificial intelligence; ChatGPT; Malnutrition; Older adults; Accuracy

INTRODUCTION

Artificial intelligence (AI) and large language models (LLMs) in particular are increasingly being used in healthcare. In healthcare delivery, the contribution of AI in enhancing clinical decision-making has garnered the most attention, particularly concerning prognostic assessment, diagnostic accuracy, treatment, clinician workflow optimization, and expansion of clinical expertise1. AI has demonstrated its potential in enhancing the quality of medical documentation and synthesizing information for clinicians. These advancements can alleviate the cognitive burden on healthcare practitioners and rectify the deficiencies inherent in human-generated documentation2. By integrating and analyzing diverse data types such as electronic health records, genomic information, and imaging results, LLMs are increasingly positioned to support clinical decision-making in treatment selection3. An increasing number of studies are being conducted on the use of AI in various fields of geriatric medicine. A study evaluating responses provided by AI to questions related to geriatric medicine revealed that scores varied significantly depending on the area of knowledge, with responses concerning the diagnosis/performing of complementary tests receiving the lowest scores4.

Malnutrition is an increasingly recognized geriatric syndrome associated with morbidity, mortality, and increased costs of care. Improving early recognition and treatment of malnutrition is crucial. Many cases of malnutrition go unnoticed, which leads to further complications and increases mortality. A meta-analysis has shown that the prevalence of malnutrition and the risk of malnutrition in older adults with dementia reach approximately 80%5. In certain conditions commonly seen in older adults, nutritional assessment should be prioritized. In cases with hip fractures and pressure ulcers, adequate and early nutritional support is crucial to promote rapid recovery, prevent complications and ensure independence, even if the patient is not malnourished6,7.

A recent review has shown that AI algorithms could identify new relationships between diet and disease outcomes, enabling clinicians to make evidence-based nutrition recommendations8. In another study, researchers developed the Malnutrition Universal Screening Tool (MUST)-Plus, a machine learning–based screening tool, which markedly enhanced the early identification and documentation of malnutrition and was well accepted by registered dietitians9. ChatGPT has strong capabilities in nutritional evaluation by estimating caloric requirements and recommending nutrient-dense foods, as well as in identifying nutritional issues using technical terminology10.

One of the primary roles of AI is to enhance clinical decision-making and streamline the delivery of healthcare services11. However, there is little evidence on the quality, accuracy, applicability, ethical challenges, and safety of these applications12. Therefore, we assessed the extent to which ChatGPT provided accurate nutritional advice among older adults in the present study.

METHODS

The study was conducted on 29 June 2025. First, twelve questions regarding general information, diagnosis, treatment, and follow-up related to malnutrition were identified and input into ChatGPT-4o mini. For each question, ChatGPT was asked to respond like a healthcare professional and provide references. To ensure that the results were not affected by previous queries, the browsing history data was completely deleted before each question (Figure 1). Ethics committee approval was not required, as patient data was not used in the present study.

Figure 1
Clinical scenarios and questions related to malnutrition and clinical scenarios.

Clinical scenarios

Three clinical scenarios mimicking real patient information were created (Figure 1).

Scenario 1: A 72-year-old female patient with decreased food intake, difficulty swallowing, advanced dementia, malnutrition, sarcopenia, and clinical findings consistent with aspiration pneumonia.

Scenario 2: A 72-year-old female patient with no chronic illness or medication use who developed pressure ulcers after a long stay in the intensive care unit, with no decrease in food intake and difficulty swallowing.

Scenario 3: A 72-year-old female patient with type 2 diabetes mellitus and osteoporosis who underwent surgery for a hip fracture, with no decrease in food intake and difficulty swallowing.

The scenarios were uploaded to ChatGPT, and the same set of 11 questions was asked to ChatGPT for each scenario (Figure 1). ChatGPT was asked to respond like a healthcare professional to each set of questions and provide references for each answer. Browsing history data was completely deleted before each question set entry.

Assessment of the responses

ChatGPT's responses were independently examined by three geriatricians. The quality of the responses was assessed using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool. It consists of six items: accuracy, clarity, relevance, completeness, sources, and usefulness. Each item has a score between 1 (strongly disagree) and 5 (strongly agree), and the sum of the scores yields a total QAMAI score between 6 and 30. The quality grades of the responses are classified as follows based on the total QAMAI score: excellent quality (30 points), very good quality (24–29 points), good quality (18–23 points), fair quality (12–17 points), and poor quality (6–11 points)13.

Statistical analysis

The inter-rater reliability among the three authors was calculated using the intraclass correlation coefficient (ICC). An excellent inter-rater reliability with an ICC of 0.84 (95%CI 0.77–0.89; p<0.001) was calculated. The normality of the data was assessed using the Shapiro-Wilk test. Continuous data are represented as mean±standard deviation. The mean QAMAI scores of the responses generated by ChatGPT were evaluated using the independent samples t-test. The significance level was 0.05, with a confidence interval of 95%. Statistical Package for the Social Sciences (SPSS) for Windows version 22.0 package program was used for statistical analysis.

RESULTS

The total mean QAMAI score for the ChatGPT-generated responses to questions related to malnutrition was 26.60, indicating very good quality. ChatGPT scored the lowest in the use of sources across all items. In terms of use of sources, responses related to general information and diagnosis had statistically significantly lower scores than responses related to treatment/follow-up (Table 1).

Table 1
The Quality Analysis of Medical Artificial Intelligence scores for responses generated by ChatGPT to questions related to malnutrition.

Some of ChatGPT's answers had obvious shortcomings. For example, when asked which tools/tests are used to diagnose malnutrition and determine its severity, it did not mention the Nutritional Risk Screening 2002 (NRS-2002) tool, which is recommended for screening hospitalized patients. In response to the question regarding nutritional needs in older individuals, it was not specified that calorie requirements should be determined based on body weight and disease status. Furthermore, it was not mentioned that protein requirements may increase in the presence of acute and chronic diseases.

In the evaluation of responses to the clinical scenarios, the lowest scores were found to be those related to the use of sources. The total QAMAI score for the clinical scenario involving the patient with a hip fracture was statistically significantly lower than the other scenarios. Accuracy, relevance, and use of sources scores were also statistically significantly lower (Table 2). ChatGPT stated that nutritional support treatment was not immediately necessary for the patient with a hip fracture.

Table 2
Comparison of the Quality Analysis of Medical Artificial Intelligence scores of ChatGPT-generated responses to the questions related to clinical scenarios.

DISCUSSION

This study highlights the importance of collaboration between AI technology and healthcare professionals to assess the accuracy of ChatGPT in cases of malnutrition. The overall performance of ChatGPT can be described as very good, as it was able to generate accurate and elaborate responses. A notable limitation was the absence of nutritional support recommendations for specific conditions, along with the lack of adjustment of calorie and protein requirements to disease status and body weight, and the insufficient use of appropriate references.

A recent review evaluating AI and clinical nutrition proposed assessing the clinical efficacy and safety of AI-powered nutrition interventions8. AI applications have been developed in many medical specialities, and positive results have been reported. They have the potential to be adapted to clinical nutrition. Nevertheless, it remains essential to ensure that the algorithms employed are transparent, reliable, and clinically validated before they are integrated into routine practice.

ChatGPT's response to the prevalence question on malnutrition indicated that 1 to 10% of community-dwelling older adults are affected by malnutrition. However, it has been shown that the prevalence of malnutrition among older adults worldwide is 18.6%14. The lower prevalence of malnutrition reported by ChatGPT may be attributed to limitations in the training data.

When asked which tools/tests are used to diagnose malnutrition and determine its severity, ChatGPT's response was not accurate enough. Because it did not mention the NRS-2002 tool, which is recommended for screening hospitalized patients. The 2024 ESPEN guideline states that amino acids and β-hydroxy-β-methylbutyrate (HMB) can be added to oral/enteral nutrition to accelerate the healing of pressure ulcers, based on the results of randomized controlled trials15. In our study, ChatGPT did not mention the use of immunonutrition supplements (arginine, glutamine, and HMB).

Based on the results of our study, ChatGPT has been found to have shortcomings in nutritional support treatment for specific conditions, such as hip fractures. It was stated that nutritional support treatment was not immediately necessary for the patient with a hip fracture. Studies and evidence-based guidelines support the use of oral nutritional supplements (ONSs) for all older patients with hip fractures to reduce nutrition-related complications and enhance clinical outcomes16,17. ChatGPT's response may be due to its access to free and publicly available sources, but not to paid, up-to-date sources and guidelines.

In our study, ChatGPT did not specify that calorie and protein requirements should be determined based on body weight and disease status. The reference values for energy intake in older adults are approximately 30 kcal and a minimum of 1 g of protein per kilogram body weight per day, which should be adjusted according to nutritional status, disease status, level of physical activity and tolerability17. Additionally, elevated nutritional demands such as those associated with muscle development, tissue repair in cases of malnutrition or wound healing, or heightened metabolic needs during illness should be addressed through a corresponding increase in dietary intake18.

The findings of this study reveal that ChatGPT scored the lowest in the use of sources across all items. Although the lack of appropriate references is a limitation, it should be emphasized that a significant portion of ChatGPT's suggestions are consistent with current clinical practice.

A recent study found that integrating an AI-based rapid nutrition diagnostic system into routine in-hospital care significantly improved recovery rates and demonstrated high cost-effectiveness19. Machine learning algorithms, by incorporating diverse factors such as age, underlying disease etiology, comorbid conditions, and laboratory parameters, are capable of providing accurate survival predictions while also estimating the potential for improvement in nutritional status. The integration of AI chatbots into routine inpatient care for nutritional deficiencies can make a significant contribution to the medical profession in treating cases of malnutrition.

Our study has some limitations. First, ChatGPT's responses have been evaluated by a relatively small number of experts. Including experts from different fields with experience in nutrition could have increased the assessment's diversity. Second, chatbots other than ChatGPT could also be included to assess their competence in the area of malnutrition and its management. Thirdly, although patients’ body mass index data were available in clinical scenarios, there was no information regarding the amount and duration of weight loss.

Despite these limitations, the study also has some strengths. The areas of knowledge covered by the set of questions were broad, and the quality of the responses was assessed using a validated tool (QAMAI). There was no information in the clinical scenarios that could lead to misunderstanding or misinterpretations, which allowed ChatGPT to be evaluated more objectively.

In conclusion, this study highlights ChatGPT's potential as a valuable tool for providing rapid and accurate information in geriatrics, particularly in the field of clinical nutrition. However, experts emphasized that ChatGPT's answers to some questions were not accurate enough. This was particularly noticeable in its failure to provide dietary recommendations for specific situations, as it did not mention the calorie and protein requirements based on weight and health conditions, and it did not provide adequate references. It is therefore of utmost importance that the responses provided by ChatGPT be critically evaluated and verified. AI-powered decision support systems could assist clinicians in making personalized treatment decisions for patients’ malnutrition.

  • Funding:
    none.
  • ETHICS APPROVAL
    Ethics committee approval was not required as patient data were not used in the present study.

DATA AVAILABILITY STATEMENT

The datasets generated and/or analyzed during the current study are available from the corresponding author upon reasonable request.

REFERENCES

  • 1 Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56. https://doi.org/10.1038/s41591-018-0300-7
    » https://doi.org/10.1038/s41591-018-0300-7
  • 2 Croxford E, Gao Y, Pellegrino N, Wong K, Wills G, First E, et al. Current and future state of evaluation of large language models for medical summarization tasks. Npj Health Syst. 2025;2:6. https://doi.org/10.1038/s44401-024-00011-2
    » https://doi.org/10.1038/s44401-024-00011-2
  • 3 Maddox TM, Embí P, Gerhart J, Goldsack J, Parikh RB, Sarich TC. Generative AI in medicine -evaluating progress and challenges. N Engl J Med. 2025;26;392(24):2479-83.
  • 4 Rosselló-Jiménez D, Docampo S, Collado Y, Cuadra-Llopart L, Riba F, Llonch-Masriera M. Geriatrics and artificial intelligence in Spain (Ger-IA project): talking to ChatGPT, a nationwide survey. Eur Geriatr Med. 2024;15(4):1129-36. https://doi.org/10.1007/s41999-024-00970-7
    » https://doi.org/10.1007/s41999-024-00970-7
  • 5 Perry E, Walton K, Lambert K. Prevalence of malnutrition in people with dementia in long-term care: a systematic review and meta-analysis. Nutrients. 2023;28;15(13):2927.
  • 6 Munoz N, Posthauer ME, Cereda E, Schols JMGA, Haesler E. The role of nutrition for pressure injury prevention and healing: the 2019 international clinical practice guideline recommendations. Adv Skin Wound Care. 2020;33(3):123-36. https://doi.org/10.1097/01.ASW.0000653144.90739.ad
    » https://doi.org/10.1097/01.ASW.0000653144.90739.ad
  • 7 Rempel AN, Rigassio Radler DL, Zelig RS. Effects of the use of oral nutrition supplements on clinical outcomes among patients who have undergone surgery for hip fracture: a literature review. Nutr Clin Pract. 2023;38(4):775-89. https://doi.org/10.1002/ncp.10980
    » https://doi.org/10.1002/ncp.10980
  • 8 Bond A, Mccay K, Lal S. Artificial intelligence & clinical nutrition: what the future might have in store. Clin Nutr ESPEN. 2023;57:542-9. https://doi.org/10.1016/j.clnesp.2023.07.082
    » https://doi.org/10.1016/j.clnesp.2023.07.082
  • 9 Parchure P, Besculides M, Zhan S, Cheng FY, Timsina P, Cheertirala SN, et al. Malnutrition risk assessment using a machine learning-based screening tool: a multicentre retrospective cohort. J Hum Nutr Diet. 2024;37(3):622-32. https://doi.org/10.1111/jhn.13286
    » https://doi.org/10.1111/jhn.13286
  • 10 Luis D. Inteligencia artificial generativa ChatGPT en nutrición clínica: avances y desafíos [Generative artificial intelligence ChatGPT in clinical nutrition - Advances and challenges]. Nutr Hosp. 2025;4;42(4):797-806. Spanish.
  • 11 Martinez-Martin N, Luo Z, Kaushal A, Adeli E, Haque A, Kelly SS, et al. Ethical issues in using ambient intelligence in health-care settings. Lancet Digit Health. 2021;3(2):e115-23. https://doi.org/10.1016/S2589-7500(20)30275-2
    » https://doi.org/10.1016/S2589-7500(20)30275-2
  • 12 Hutson M. Robo-writers: the rise and risks of language-generating AI. Nature. 2021;591(7848):22-5. https://doi.org/10.1038/d41586-021-00530-0
    » https://doi.org/10.1038/d41586-021-00530-0
  • 13 Vaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA, et al. Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol. 2024;281(11):6123-31. https://doi.org/10.1007/s00405-024-08710-0
    » https://doi.org/10.1007/s00405-024-08710-0
  • 14 Salari N, Darvishi N, Bartina Y, Keshavarzi F, Hosseinian-Far M, Mohammadi M. Global prevalence of malnutrition in older adults: a comprehensive systematic review and meta-analysis. Public Health Pract (Oxf). 2025;9:100583. https://doi.org/10.1016/j.puhip.2025.100583
    » https://doi.org/10.1016/j.puhip.2025.100583
  • 15 Wunderle C, Gomes F, Schuetz P, et al. ESPEN practical guideline: nutritional support for polymorbid medical inpatients. Clin Nutr. 2024;43(3):674-91.
  • 16 Chen B, Zhang JH, Duckworth AD, Clement ND. Effect of oral nutritional supplementation on outcomes in older adults with hip fractures and factors influencing compliance. Bone Joint J. 2023;105-B(11):1149-58. https://doi.org/10.1302/0301-620X.105B11.BJJ-2023-0139.R1
    » https://doi.org/10.1302/0301-620X.105B11.BJJ-2023-0139.R1
  • 17 Volkert D, Beck AM, Cederholm T, Cruz-Jentoft A, Hooper L, Kiesswetter E, et al. ESPEN practical guideline: clinical nutrition and hydration in geriatrics. Clin Nutr. 2022;41(4):958-89. https://doi.org/10.1016/j.clnu.2022.01.024
    » https://doi.org/10.1016/j.clnu.2022.01.024
  • 18 Cruz-Jentoft AJ, Volkert D. Malnutrition in older adults. N Engl J Med. 2025;392(22):2244-55. https://doi.org/10.1056/NEJMra2412275
    » https://doi.org/10.1056/NEJMra2412275
  • 19 Sun MY, Wang Y, Zheng T, Wang X, Lin F, Zheng L-Y, et al. Health economic evaluation of an artificial intelligence (AI)-based rapid nutritional diagnostic system for hospitalised patients: a multicentre, randomised controlled trial. Clin Nutr. 2024;43(10):2327-35.

Edited by

Publication Dates

  • Publication in this collection
    29 June 2026
  • Date of issue
    2026

History

  • Received
    17 Sept 2025
  • Accepted
    26 Mar 2026
location_on
Associação Médica Brasileira R. São Carlos do Pinhal, 324, 01333-903 São Paulo SP - Brazil, Tel: +55 11 3178-6800, Fax: +55 11 3178-6816 - São Paulo - SP - Brazil
E-mail: ramb@amb.org.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro