LLM Performance Degradation Across Pediatric Age Groups: Demographic Stratified Evaluation of Clinical NLP on Synthetic Patient Cohorts

Authors

  • Vinod Rufus Motani

Keywords:

Large Language Models, Clinical NLP, Pediatric Healthcare, Synthetic Patient Cohorts, Named Entity Recognition, Demographic Fairness, Electronic Health Records.

Abstract

Large language models (LLMs) have become widespread in the context of clinical natural language processing (NLP). However, their effectiveness in pediatric age groups requires further research. The present study seeks to analyze the decrease in LLM performance on the basis of the demographic-stratified analysis of clinical NLP tasks performed on synthetic pediatric patient cohorts. The research is based on a qualitative secondary approach with the usage of evidences from recent articles about pediatric artificial intelligence, clinical NLP, fairness of models, and synthetic health data. The cohorts in question have been divided into four age groups: neonatal, infant, child, and adolescent. It has been found out that there is a considerable difference in the performance of LLMs among different developmental periods. Adolescent records showed the best results in terms of accuracy as their documentation format resembles the documentation format of adult electronic health records. As for neonatal and infant records, the accuracy, recall, and F1-score of those were lower because of specialized language, physiological features, age-based lab references, and brief clinical notes. Demographic fairness analysis also revealed discrepancies in terms of prediction accuracy for different pediatric populations. The study also emphasizes the importance of using synthetic pediatric cohorts for conducting privacy-safe model evaluation. Synthetic data made it possible to conduct controlled benchmarking, minimize the problem of data shortage, and conduct reproducible evaluation while not revealing any patient data. Generally, the findings show that the existing LLMs perform inconsistently among pediatric populations and need to be adjusted, audited for fairness, and trained specifically․

Downloads

Download data is not yet available.

References

1. Greyling, C., 2023. How Does Large Language Models Use Long Contexts? https://www.humanfirst.ai/blog/how-does-large-language-models-use-long-contexts

2. Bose, P., Srinivasan, S., Sleeman IV, W. C., Palta, J., Kapoor, R., & Ghosh, P. (2021). A survey on recent named entity recognition and relationship extraction techniques on clinical texts. Applied Sciences, 11(18), 8319. https://doi.org/10.3390/app11188319

3. Gessesse, A. D., Belete, M. B., & Tadesse, F. (2024). Time, cause of early neonatal death, and its predictors among neonates admitted to neonatal intensive care units at Bahir Dar City public hospitals, northwest Ethiopia: A prospective follow-up study. Frontiers in Pediatrics, 12, 1335858. https://doi.org/10.3389/fped.2024.1335858

4. O'Sullivan, E., van de Lande, L. S., Oosting, A.-J. C., Papaioannou, A., Jeelani, N. O., Koudstaal, M. J., Khonsari, R. H., Dunaway, D. J., Zafeiriou, S., & Schievano, S. (2021). The 3D skull 0–4 years: A validated, generative, statistical shape model. Medical Image Analysis, 75, 102265. https://doi.org/10.1016/j.media.2021.102265

5. Schepens, J., Marx, N., & Gagl, B. (2023). Can we utilize large language models (LLMs) to generate useful linguistic corpora? A case study of the word frequency effect in young German readers. Preprint from PsyArXiv https://doi. org/10.31234/osf. io/gm9b6. https://files.osf.io/v1/resources/gm9b6_v1/providers/osfstorage/6532d1fe164d320baaa5e643?action=download&direct&version=4

6. Wornow, M., Xu, Y., Thapa, R., Patel, B., Steinberg, E., Fleming, S., ... & Shah, N. H. (2023). The shaky foundations of large language models and foundation models for electronic health records. npj digital medicine, 6(1), 135. https://www.nature.com/articles/s41746-023-00879-8

7. Raza, S., & Schwartz, B. (2023). Entity and relation extraction from clinical case reports of COVID-19: a natural language processing approach. BMC Medical Informatics and Decision Making, 23(1), 20. https://link.springer.com/article/10.1186/s12911-023-02117-3

8. Murray, L., Gopinath, D., Agrawal, M., Horng, S., Sontag, D., & Karger, D. R. (2021, October). Medknowts: Unified documentation and information retrieval for electronic health records. In The 34th Annual ACM Symposium on User Interface Software and Technology (pp. 1169-1183). https://dl.acm.org/doi/abs/10.1145/3472749.3474814

9. Mumtaz, U., Ahmed, A., & Mumtaz, S. (2023). LLMs-Healthcare: Current applications and challenges of large language models in various medical specialties. arXiv preprint arXiv:2311.12882. https://arxiv.org/abs/2311.12882

10. Vadyala, S. R., & Sherer, E. A. (2021). Natural language processing accurately categorizes indications, findings and pathology reports from multicenter colonoscopy. arXiv preprint arXiv:2108.11034. https://arxiv.org/abs/2106.14463

11. Yacouby, R., & Axman, D. (2020, November). Probabilistic extension of precision, recall, and f1 score for more thorough evaluation of classification models. In Proceedings of the first workshop on evaluation and comparison of NLP systems (pp. 79-91). https://aclanthology.org/2020.eval4nlp-1.9/

12. Tucker, A., Wang, Z., Rotalinti, Y., & Myles, P. (2020). Generating high-fidelity synthetic patient data for assessing machine learning healthcare software. NPJ digital medicine, 3(1), 147. https://www.nature.com/articles/s41746-020-00353-9

13. Shi, W., Zhuang, Y., Zhu, Y., Iwinski, H., Wattenbarger, M., & Wang, M. D. (2023, September). Retrieval-augmented large language models for adolescent idiopathic scoliosis patients in shared decision-making. In Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics (pp. 1-10). https://dl.acm.org/doi/abs/10.1145/3584371.3612956

14. Liu, J., Gallego, B., & Barbieri, S. (2022). Incorporating uncertainty in learning to defer algorithms for safe computer-aided diagnosis. Scientific reports, 12(1), 1762. https://www.nature.com/articles/s41598-022-05725-7

15. Keles, E., & Bagci, U. (2023). The past, current, and future of neonatal intensive care units with artificial intelligence: a systematic review. NPJ digital medicine, 6(1), 220. https://www.nature.com/articles/s41746-023-00941-5

16. Day, L. T., Gore-Langton, G. R., Rahman, A. E., Basnet, O., Shabani, J., Tahsina, T., ... & Lawn, J. E. (2020). Labour and delivery ward register data availability, quality, and utility-Every Newborn-birth indicators research tracking in hospitals (EN-BIRTH) study baseline analysis in three countries. BMC health services research, 20(1), 737. https://link.springer.com/article/10.1186/s12913-020-5028-7

17. Bhanot, K. (2023). Synthetic data generation and evaluation for fairness (Doctoral dissertation, Rensselaer Polytechnic Institute). https://search.proquest.com/openview/56fc52700b150cc4920289ab6be0865b/1?pq-origsite=gscholar&cbl=18750&diss=y

18. Chen, R. J., Wang, J. J., Williamson, D. F., Chen, T. Y., Lipkova, J., Lu, M. Y., ... & Mahmood, F. (2023). Algorithmic fairness in artificial intelligence for medicine and healthcare. Nature biomedical engineering, 7(6), 719-742. https://www.nature.com/articles/s41551-023-01056-8

19. Sacco, S. J., Chen, K., Wang, F., & Aseltine, R. (2023). Target-based fusion using social determinants of health to enhance suicide prediction with electronic health records. PloS one, 18(4), e0283595. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0283595.

Downloads

Published

2025-11-21

How to Cite

1.
Motani VR. LLM Performance Degradation Across Pediatric Age Groups: Demographic Stratified Evaluation of Clinical NLP on Synthetic Patient Cohorts. J Neonatal Surg [Internet]. 2025 Nov. 21 [cited 2026 Oct. 6];14(32S):11250-6. Available from: https://jneonatalsurg.com/index.php/jns/article/view/10606