SUMMARY
OBJECTIVE: The aim of the study was to evaluate the quality and readability of ChatGPT responses to frequently asked questions by individuals with posture disorder. Providing reliable and evidence-based information about posture disorders is vital for individuals to be correctly informed.
METHODS: The 10 most frequently asked questions about posture disorder were selected by two researchers from a list created by ChatGPT. The questions were transmitted to ChatGPT 4.0, and the initial responses were recorded without further follow-up questions. The quality of the responses was then assessed by five independent experts (three physiotherapists, one physical therapy and rehabilitation specialist, and one orthopedics and traumatology specialist) with a four-grade evaluation system. Readability levels were analyzed with the Flesch-Kincaid Grade Level through WordCalc software. Statistical analysis was performed using Statistical Package for the Social Sciences v29.0, and intraclass correlation coefficients were used to measure inter-rater reliability.
RESULTS: Following a thorough evaluation of the 10 responses received, six were rated as "Excellent responses requiring no explanation," while a further four were designated as "Satisfactory responses requiring minimal explanation." The median quality score of the responses was high, indicating good alignment with current evidence-based practice. The average readability level of the responses was determined to be 8.4. Inter-rater reliability was good, with an intraclass correlation coefficients value of 0.756.
CONCLUSION: ChatGPT provides relatively coherent and generally readable answers to frequently asked questions about posture disorders, with most needing minimal explanation. While promising as a resource to meet the information needs of people with posture disorders, further improvements are needed to align it with personalized health needs.
KEYWORDS:
Artificial intelligence; Physical therapy; Posture
INTRODUCTION
Recent years have witnessed significant advancements in the field of artificial intelligence (AI) technology. AI-supported chatbots are computer programs capable of understanding human language and generating appropriate responses, and can interact with users in depth1. These programs are trained with deep learning algorithms over large amounts of text data from the internet, enabling them to develop an understanding of and respond to a wide range of topics2. ChatGPT, an AI chatbot developed by OpenAI, employs deep learning technology to provide human-like responses to natural language inputs3. ChatGPT has attracted attention for its user-friendly interface and has the potential to create substantial impacts on healthcare services and patient education4.
Posture is defined as the position of the body when sitting or standing5. Variations in posture may be attributable to age, gender, body development, and environmental or psychological factors6. Postural disorder (PD) is a multifactorial condition that arises from muscle imbalances, nutritional issues, inactivity, poor posture habits, improper bag carrying, and psychological factors7. PD can lead to muscle weakness, resulting in aesthetic problems, pain, and musculoskeletal diseases8. Additionally, PD is a prevalent yet often disregarded health concern on a global scale9. Therefore, individuals affected by PDs may turn to AI systems like ChatGPT for the procurement of information.
Despite its current limitations, ChatGPT demonstrates considerable potential in filtering scientific content and contributing to the development of personalized treatment processes10. A recent study11 drew parallels between ChatGPT's recommendations for various musculoskeletal disorders and the APTA Academy of Orthopedic Physical Therapy's Clinical Practice Guidelines (CPGs), finding that ChatGPT exhibited high levels of accuracy and reliability. Researchers have indicated that ChatGPT has potential as a clinical support tool in rehabilitation. Furthermore, several other studies evaluated ChatGPT's responses to the most frequently asked questions from patients about hip arthroscopy12, ankle arthroplasty13, and total shoulder arthroplasty14. In these studies, ChatGPT's responses to the aforementioned questions were found to be generally satisfactory and adequate, with minimal additional clarification required in most cases.
The quality of the responses provided by ChatGPT to questions regarding PD remains to be ascertained. The objective of this study was to evaluate the quality and readability of the responses provided by ChatGPT to inquiries concerning PD.
METHODS
Question collection process
A prompt was entered into the ChatGPT application as "What are the 50 most frequently asked questions by those with PDs?" Following a thorough examination of the pool of questions, a total of 10 questions were selected by two researchers (two physiotherapists) and can be seen in Table 1.
ChatGPT usage
The questions were formulated using ChatGPT 4.0. In order to prevent ChatGPT from being influenced by prior interactions, the browser history was initially erased, and then a new ChatGPT account was established for the initial use. In the subsequent phase, 10 questions determined by the researchers were asked in order, and the initial responses of ChatGPT to each question were meticulously recorded. At this time, no further questions were posed, and no additional explanations were made.
Evaluation
The quality of the responses was assessed by five independent researchers (three physiotherapists, one physical therapy and rehabilitation specialist, and one orthopedics and traumatology specialist) using the four-grade system proposed by Mika et al.15:
-
"Excellent response, no explanation needed": The answer was considered extremely accurate and comprehensive, thus providing information without requiring any additional explanation.
-
"Satisfactory, requires minimal explanation": The answer was rated correct but required minimal supplementary explanation to fully address the user's question.
-
"Satisfactory, needs moderate explanation": Although the answer was correct, moderate additional explanation was necessary to meet the user's needs.
-
"Unsatisfactory, requires significant explanation": The response contains significant misinformation or is considered overly generalized, which may lead to user misunderstanding.
The readability of ChatGPT responses was assessed by means of WordCalc software, whereby responses to each question were pasted into a readability calculator, and the corresponding Flesch-Kincaid Rating Level was recorded.
Statistics
The data evaluation process involved the utilization of the SPSS (Statistical Package for the Social Sciences) Statistics for Windows, Version 29.0 (IBM Corp., Armonk, NY, USA) software package. The results of the quality assessment of the responses are given as median (min–max). The inter-rater reliability was calculated using intraclass correlation coefficients (ICC). The ICC values were considered and then categorized as poor for <0.50, moderate for 0.50–0.75, good for 0.75–0.90, and excellent for >0.9016.
RESULTS
The study was conducted with ChatGPT version 4.0 in February 2025. Six of the 10 responses evaluated in the study were evaluated and classified as "Excellent response, no explanation needed," while four were designated as "Satisfactory, requires minimal explanation." In addition to the median scores of the responses, the average readability level was determined to be 8.4, as shown in Table 2. Differences among expert evaluations were observed for certain questions, reflecting a diversity of opinions among the raters; this is illustrated in Table 3.
Question 1: The question "Which doctor should I see to treat a posture disorder?" was evaluated as "Excellent response, no explanation needed" by three evaluators, as "Satisfactory, requires minimal explanation" by one evaluator, and as "Satisfactory, needs moderate explanation" by one evaluator, and the readability level was determined as 12.
Question 2: The question "How do I know whether I have a posture disorder?" was evaluated as "Excellent response, no explanation needed" by three evaluators and as "Satisfactory, requires minimal explanation" by two evaluators, and the readability level was calculated as seven.
Question 3: The question "I'm constantly hunched over; how can I fix it?" was evaluated by two evaluators as "Excellent response, no explanation needed" and by three evaluators as "Satisfactory, requires minimal explanation," and the readability level was ascertained as 6.3.
Question 4: The question "If I use the wrong pillow and mattress, will my posture be damaged?" was assessed by all evaluators as "Excellent response, no explanation needed," and the readability level was established as 6.7.
Question 5: For the question "Which exercises can I do to improve my posture?," the readability level was determined to be 5.4, with one assessor rating the answer as "Excellent response, no explanation needed" and four assessors rating the answer as "Satisfactory, requires minimal explanation."
Question 6: The readability level of the question "How long does it take to correct my posture?" was 6.7, and one assessor rated the answer as "Excellent response, no explanation needed," two assessors rated it as "Satisfactory, requires minimal explanation," and two assessors rated it as "Satisfactory, needs moderate explanation."
Question 7: The question "Should I use a posture correcting corset?" was evaluated as "Excellent response, no explanation needed" by all evaluators, and the readability level was determined as 8.3.
Question 8: The question "Can posture disorders that occur in childhood be corrected in adulthood?" was evaluated by two evaluators as "Excellent response, no explanation needed" and three evaluators as "Satisfactory, requires minimal explanation," and the readability level was established as 12.
Question 9: The question "Does excess weight impair my posture?" was evaluated as "Excellent response, no explanation needed" by three assessors and as "Satisfactory, requires minimal explanation" by two assessors, and the readability level was found to be 10.2.
Question 10: The readability level of the question "Does carrying a bag on one shoulder ruin my posture?" scored 9.4, with four assessors rating the answer as "Excellent response, no explanation needed" and one assessor rating it as "Satisfactory, requires minimal explanation."
DISCUSSION
In this study, the quality and readability of ChatGPT 4.0 responses related to PD were evaluated. ChatGPT provides answers to frequently asked questions about PDs that are partially consistent with expert opinions, but further verification is required to confirm complete clinical accuracy. In terms of readability, the responses were found to be above the recommended sixth-grade level for patient education materials. The average Flesch-Kincaid reading level was calculated as 8.4, which is above the recommended level for general patient-focused content. This suggests that ChatGPT responses may negatively impact their understandability, particularly for individuals with limited health literacy.
This study is pioneering, namely the first study, in its evaluation of ChatGPT responses on PDs. In the physiotherapy and musculoskeletal system, the accuracy and effectiveness of ChatGPT have been examined through various studies and methodologies, yielding different results. One recent study found that ChatGPT answered clinical questions related to the musculoskeletal system with 80% accuracy, and its responses were consistent in terms of reliability11. These findings serve to reinforce the potential of ChatGPT to function as a reliable resource within the domain of clinical decision support systems. AlShehri et al. assessed the responses to common patient questions about hip arthroscopy and noted that ChatGPT 3.5 provided satisfactory findings in terms of accuracy and completeness, though some errors should be considered before using it for patient education17. Similarly, the research by Artioli et al. found that most ChatGPT responses regarding ankle arthroplasty were rated as excellent, although in some instances, more detailed explanations were needed13. Conversely, the result of another study revealed that ChatGPT offered effective evidence-based answers to frequently asked patient questions in evaluations regarding total hip arthroplasty15. Assessments of the accuracy and consistency of ChatGPT's responses to inquiries related to total knee arthroplasty18,19 indicate that this system holds significant potential as a patient education and clinical decision support tool.
Our study and similar studies in the literature demonstrate the potential of ChatGPT to provide accurate and effective information to patients regarding the musculoskeletal system. However, in some cases, more detailed explanations and up-to-date information are required, which may limit the use of ChatGPT as a clinical decision support tool. There is a need for further research into the ability of ChatGPT to provide accurate and reliable information to patients.
In order to enhance the objectivity of ChatGPT's responses and ensure that the evaluations are grounded in evidence-based principles, an additional comparative analysis was conducted. Within this scope, the responses provided by ChatGPT to the selected questions were compared with peer-reviewed academic sources. Table 4 presents this comparison, highlighting both areas where the responses closely align with scientific literature and instances where content limitations were identified. This approach was adopted to reduce the subjectivity inherent in expert-based evaluations and to enable a more standardized and reliable framework for assessing the accuracy of AI-generated health information.
Strengths and limitations
This study has several strengths that increase its relevance for assessing the role of ChatGPT in PD. Firstly, this study is the first to assess the quality and accuracy of AI responses in relation to PD. This focus meets a significant knowledge gap by handling the need for accurate and accessible information about PD and related treatment methods.
This study has several limitations. Although this study aimed to identify frequently asked questions (FAQs) about PDs, the selected questions were generated by ChatGPT. This introduces a potential selection bias, as the questions may reflect the assumptions of the AI rather than the actual needs of patients. The answers provided by ChatGPT can sometimes be superficial for complex and multidisciplinary health issues, such as posture problems. Moreover, some of the included questions are inherently ambiguous, and their answers may vary significantly depending on individual factors. Given the complex and personalized nature of PDs, the responses provided by ChatGPT are often limited to generalized information. The model is also prone to generating hallucinations, such as fabricated references or inaccurate clinical suggestions. Furthermore, since ChatGPT has been trained predominantly on data from Western medical perspectives, its responses may lack relevance or sensitivity to different cultural and regional contexts. Especially in situations that require clinical decisions, more in-depth and expert-based information is needed. Additionally, the process of relying on expert judgment to assess the quality of responses can introduce subjectivity, even when standardized criteria are applied. Another limitation is that ChatGPT is not connected to a constantly updated source database. Therefore, there is a risk that future research on PD may not be reflected in the model. It should also be acknowledged that AI models are continuously evolving, and future versions may demonstrate significantly different performance. Therefore, comparative evaluations across different AI versions are essential. In our study, the reading level of the content was found to be above ideal standards. This may pose accessibility challenges for individuals with low health literacy and highlights the importance of optimizing readability in the future development of AI-generated health content. Future research should investigate the accuracy and effectiveness of the information provided by ChatGPT regarding PDs with different user profiles and evaluate the management of these conditions using multidisciplinary approaches.
CONCLUSION
ChatGPT is a promising complementary tool for providing education and information about PD, but it cannot replace professional health advice. Developments in AI technology and regular refinement in line with updated scientific evidence may increase the reliability and effectiveness of the model. Future research should monitor the quality of the model's responses over time in order to evaluate performance trends such as consistency, improvement, or decline. The impact of the information provided by ChatGPT on clinical outcomes—particularly in terms of patient adherence, comprehension, and behavioral change—should be examined in detail. Moreover, the role of multi-step and interactive dialogues in real-life patient education may constitute a distinct area of investigation. In addition, conducting version-based comparisons among continuously evolving AI models is essential to identify potential differences in performance. Future studies should aim to enhance ChatGPT's capacity to provide information on PDs by developing user-centered FAQ content based on data obtained from clinical interviews, surveys, or digital health platforms. In doing so, the integration of AI-powered tools into clinical practice can be grounded on a more robust foundation.
DATA AVAILABILITY STATEMENT
The datasets generated and/or analyzed during the current study are available from the corresponding author upon reasonable request.
REFERENCES
-
1 Massey PA, Montgomery C, Zhang AS. Comparison of ChatGPT-3.5, ChatGPT-4, and orthopaedic resident performance on orthopaedic assessment examinations. J Am Acad Orthop Surg. 2023;31(23):1173-9. https://doi.org/10.5435/JAAOS-D-23-00396
» https://doi.org/10.5435/JAAOS-D-23-00396 -
2 Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, et al. The future landscape of large language models in medicine. Commun Med (Lond). 2023;3(1):141. https://doi.org/10.1038/s43856-023-00370-1
» https://doi.org/10.1038/s43856-023-00370-1 -
3 Cotton DRE, Cotton PA, Shipway JR. Chatting and cheating: ensuring academic integrity in the era of ChatGPT. Innov Educ Teach Int. 2024;61(2):228-39. https://doi.org/10.1080/14703297.2023.2190148
» https://doi.org/10.1080/14703297.2023.2190148 -
4 Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. 2023;6:1169595. https://doi.org/10.3389/frai.2023.1169595
» https://doi.org/10.3389/frai.2023.1169595 -
5 Roque GC, Oliveira RG, Sorzi MV, Oliveira LC. Are pilates exercises effective in improving postural misalignment? Systematic Review and Metanalysis. Musculoskeletal Care. 2024;22(4):e70009. https://doi.org/10.1002/msc.70009
» https://doi.org/10.1002/msc.70009 -
6 Li F, Omar Dev RD, Soh KG, Wang C, Yuan Y. Effects of pilates on body posture: a systematic review. Arch Rehabil Res Clin Transl. 2024;6(3):100345. https://doi.org/10.1016/j.arrct.2024.100345
» https://doi.org/10.1016/j.arrct.2024.100345 -
7 Batebi M, Namin BG, Nasermelli MH, Abolhasani M, Fard AHS. The relationship between static and dynamic postural deformities with pain and quality of life in non-athletic women. BMC Musculoskelet Disord. 2024;25(1):771. https://doi.org/10.1186/s12891-024-07880-6
» https://doi.org/10.1186/s12891-024-07880-6 -
8 Porto AB, Nascimento Guimarães A, Alves Okazaki VH. The effect of exercise on postural alignment: a systematic review. J Bodyw Mov Ther. 2024;40:99-108. https://doi.org/10.1016/j.jbmt.2024.04.004
» https://doi.org/10.1016/j.jbmt.2024.04.004 -
9 Kasović M, Štefan L, Piler P, Zvonar M. Longitudinal associations between sport participation and fat mass with body posture in children: a 5-year follow-up from the Czech ELSPAC study. PLoS One. 2022;17(4):e0266903. https://doi.org/10.1371/journal.pone.0266903
» https://doi.org/10.1371/journal.pone.0266903 -
10 Cavazzotto TG, Dantas DB, Queiroga MR. ChatGPT and exercise prescription: human vs. machine or human plus machine? J Sport Health Sci. 2023;13(5):661-2. https://doi.org/10.1016/j.jshs.2023.10.008
» https://doi.org/10.1016/j.jshs.2023.10.008 -
11 Hao J, Yao Z, Tang Y, Remis A, Wu K, Yu X. Artificial intelligence in physical therapy: evaluating ChatGPT's role in clinical decision support for musculoskeletal care. Ann Biomed Eng. 2025;53(1):9-13. https://doi.org/10.1007/s10439-025-03676-4
» https://doi.org/10.1007/s10439-025-03676-4 -
12 Özbek EA, Ertan MB, Kından P, Karaca MO, Gürsoy S, Chahla J. ChatGPT can offer at least satisfactory responses to common patient questions regarding hip arthroscopy. Arthroscopy. 2025;41(6):1806-27. https://doi.org/10.1016/j.arthro.2024.08.036
» https://doi.org/10.1016/j.arthro.2024.08.036 -
13 Artioli E, Veronesi F, Mazzotti A, Brogini S, Zielli SO, Giavaresi G, et al. Assessing ChatGPT responses to common patient questions regarding total ankle arthroplasty. J Exp Orthop. 2024;12(1):e70138. https://doi.org/10.1002/jeo2.70138
» https://doi.org/10.1002/jeo2.70138 -
14 White CA, Masturov YA, Haunschild E, Michaelson E, Shukla DR, Cagle PJ. Can ChatGPT reliably answer the most common patient questions regarding total shoulder arthroplasty? J Shoulder Elbow Surg. 2025;34(5):e254-64. https://doi.org/10.1016/j.jse.2024.08.025
» https://doi.org/10.1016/j.jse.2024.08.025 -
15 Mika AP, Martin JR, Engstrom SM, Polkowski GG, Wilson JM. Assessing ChatGPT responses to common patient questions regarding total hip arthroplasty. J Bone Joint Surg Am. 2023;105(19):1519-26. https://doi.org/10.2106/JBJS.23.00209
» https://doi.org/10.2106/JBJS.23.00209 -
16 Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155-63. https://doi.org/10.1016/j.jcm.2016.02.012
» https://doi.org/10.1016/j.jcm.2016.02.012 -
17 AlShehri Y, McConkey M, Lodhia P. ChatGPT provides satisfactory but occasionally inaccurate answers to common patient hip arthroscopy questions. Arthroscopy. 2025;41(5):1337-47. https://doi.org/10.1016/j.arthro.2024.06.017
» https://doi.org/10.1016/j.arthro.2024.06.017 -
18 Magruder ML, Rodriguez AN, Wong JCJ, Erez O, Piuzzi NS, Scuderi GR, et al. Assessing ability for ChatGPT to answer total knee arthroplasty-related questions. J Arthroplasty. 2024;39(8):2022-7. https://doi.org/10.1016/j.arth.2024.02.023
» https://doi.org/10.1016/j.arth.2024.02.023 -
19 Zhang S, Liau ZQG, Tan KLM, Chua WL. Evaluating the accuracy and relevance of ChatGPT responses to frequently asked questions regarding total knee replacement. Knee Surg Relat Res. 2024;36(1):15. https://doi.org/10.1186/s43019-024-00218-5
» https://doi.org/10.1186/s43019-024-00218-5 -
20 Saggini R, Anastasi GP, Battilomo S, Maietta Latessa P, Costanzo G, Carlo F, et al. Consensus paper on postural dysfunction: recommendations for prevention, diagnosis and therapy. J Biol Regul Homeost Agents. 2021;35(2):441-56. https://doi.org/10.23812/20-743-A
» https://doi.org/10.23812/20-743-A -
21 Radwan A, Ashton N, Gates T, Kilmer A, VanFleet M, Alnajar M, et al. Effect of different pillow designs on promoting sleep comfort, quality, & spinal alignment: a systematic review. Eur J Integr Med. 2021;42:101269. https://doi.org/10.1016/j.eujim.2020.101269
» https://doi.org/10.1016/j.eujim.2020.101269 - 22 Yamak B, İmamoğlu O, İslamoğlu İ, Çebi M, Kaya E, Demir A, et al. The effects of exercise on body posture. Turk Stud Soc Sci. 2018;13(18):1377-88.
-
23 Bayattork M, Sköld MB, Sundstrup E, Andersen LL. Exercise interventions to improve postural malalignments in head, neck, and trunk among adolescents, adults, and older people: systematic review of randomized controlled trials. J Exerc Rehabil. 2020;16(1):36-48. https://doi.org/10.12965/jer.2040034.017
» https://doi.org/10.12965/jer.2040034.017 - 24 Lyu S. Posture modification effects using soft materials structures [dissertation]. Minneapolis (MN): University of Minnesota; 2016.
-
25 Bayartai ME, Luomajoki H, Tringali G, Micheli R, Abbruzzese L, Sartorio A. Differences in spinal posture and mobility between adults with obesity and normal weight individuals. Sci Rep. 2023;13(1):13409. https://doi.org/10.1038/s41598-023-40470-5
» https://doi.org/10.1038/s41598-023-40470-5 -
26 Rajfur JB, Rajfur KJ, Roden N, Fras-Łabanc B, Dolibog P. Evaluation of the effect of the carried baggage on the selected stabilometric parameters, body posture and occurrence of pain in young women. Med Sci Pulse. 2023;17(3):1-11. https://doi.org/10.5604/01.3001.0054.1838
» https://doi.org/10.5604/01.3001.0054.1838
Edited by
-
Scientifıc Editor:
José Maria Soares Júnior http://orcid.org/0000-0003-0774-9404
