LLM Performance Degradation Across Pediatric Age Groups: Demographic Stratified Evaluation of Clinical NLP on Synthetic Patient Cohorts
Keywords:
Large Language Models, Clinical NLP, Pediatric Healthcare, Synthetic Patient Cohorts, Named Entity Recognition, Demographic Fairness, Electronic Health Records.Abstract
Large language models (LLMs) have become widespread in the context of clinical natural language processing (NLP). However, their effectiveness in pediatric age groups requires further research. The present study seeks to analyze the decrease in LLM performance on the basis of the demographic-stratified analysis of clinical NLP tasks performed on synthetic pediatric patient cohorts. The research is based on a qualitative secondary approach with the usage of evidences from recent articles about pediatric artificial intelligence, clinical NLP, fairness of models, and synthetic health data. The cohorts in question have been divided into four age groups: neonatal, infant, child, and adolescent. It has been found out that there is a considerable difference in the performance of LLMs among different developmental periods. Adolescent records showed the best results in terms of accuracy as their documentation format resembles the documentation format of adult electronic health records. As for neonatal and infant records, the accuracy, recall, and F1-score of those were lower because of specialized language, physiological features, age-based lab references, and brief clinical notes. Demographic fairness analysis also revealed discrepancies in terms of prediction accuracy for different pediatric populations. The study also emphasizes the importance of using synthetic pediatric cohorts for conducting privacy-safe model evaluation. Synthetic data made it possible to conduct controlled benchmarking, minimize the problem of data shortage, and conduct reproducible evaluation while not revealing any patient data. Generally, the findings show that the existing LLMs perform inconsistently among pediatric populations and need to be adjusted, audited for fairness, and trained specifically․
Downloads
References
1. Greyling, C., 2023. How Does Large Language Models Use Long Contexts? https://www.humanfirst.ai/blog/how-does-large-language-models-use-long-contexts
2. Bose, P., Srinivasan, S., Sleeman IV, W. C., Palta, J., Kapoor, R., & Ghosh, P. (2021). A survey on recent named entity recognition and relationship extraction techniques on clinical texts. Applied Sciences, 11(18), 8319. https://doi.org/10.3390/app11188319
3. Gessesse, A. D., Belete, M. B., & Tadesse, F. (2024). Time, cause of early neonatal death, and its predictors among neonates admitted to neonatal intensive care units at Bahir Dar City public hospitals, northwest Ethiopia: A prospective follow-up study. Frontiers in Pediatrics, 12, 1335858. https://doi.org/10.3389/fped.2024.1335858
4. O'Sullivan, E., van de Lande, L. S., Oosting, A.-J. C., Papaioannou, A., Jeelani, N. O., Koudstaal, M. J., Khonsari, R. H., Dunaway, D. J., Zafeiriou, S., & Schievano, S. (2021). The 3D skull 0–4 years: A validated, generative, statistical shape model. Medical Image Analysis, 75, 102265. https://doi.org/10.1016/j.media.2021.102265
5. Schepens, J., Marx, N., & Gagl, B. (2023). Can we utilize large language models (LLMs) to generate useful linguistic corpora? A case study of the word frequency effect in young German readers. Preprint from PsyArXiv https://doi. org/10.31234/osf. io/gm9b6. https://files.osf.io/v1/resources/gm9b6_v1/providers/osfstorage/6532d1fe164d320baaa5e643?action=download&direct&version=4
6. Wornow, M., Xu, Y., Thapa, R., Patel, B., Steinberg, E., Fleming, S., ... & Shah, N. H. (2023). The shaky foundations of large language models and foundation models for electronic health records. npj digital medicine, 6(1), 135. https://www.nature.com/articles/s41746-023-00879-8
7. Raza, S., & Schwartz, B. (2023). Entity and relation extraction from clinical case reports of COVID-19: a natural language processing approach. BMC Medical Informatics and Decision Making, 23(1), 20. https://link.springer.com/article/10.1186/s12911-023-02117-3
8. Murray, L., Gopinath, D., Agrawal, M., Horng, S., Sontag, D., & Karger, D. R. (2021, October). Medknowts: Unified documentation and information retrieval for electronic health records. In The 34th Annual ACM Symposium on User Interface Software and Technology (pp. 1169-1183). https://dl.acm.org/doi/abs/10.1145/3472749.3474814
9. Mumtaz, U., Ahmed, A., & Mumtaz, S. (2023). LLMs-Healthcare: Current applications and challenges of large language models in various medical specialties. arXiv preprint arXiv:2311.12882. https://arxiv.org/abs/2311.12882
10. Vadyala, S. R., & Sherer, E. A. (2021). Natural language processing accurately categorizes indications, findings and pathology reports from multicenter colonoscopy. arXiv preprint arXiv:2108.11034. https://arxiv.org/abs/2106.14463
11. Yacouby, R., & Axman, D. (2020, November). Probabilistic extension of precision, recall, and f1 score for more thorough evaluation of classification models. In Proceedings of the first workshop on evaluation and comparison of NLP systems (pp. 79-91). https://aclanthology.org/2020.eval4nlp-1.9/
12. Tucker, A., Wang, Z., Rotalinti, Y., & Myles, P. (2020). Generating high-fidelity synthetic patient data for assessing machine learning healthcare software. NPJ digital medicine, 3(1), 147. https://www.nature.com/articles/s41746-020-00353-9
13. Shi, W., Zhuang, Y., Zhu, Y., Iwinski, H., Wattenbarger, M., & Wang, M. D. (2023, September). Retrieval-augmented large language models for adolescent idiopathic scoliosis patients in shared decision-making. In Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics (pp. 1-10). https://dl.acm.org/doi/abs/10.1145/3584371.3612956
14. Liu, J., Gallego, B., & Barbieri, S. (2022). Incorporating uncertainty in learning to defer algorithms for safe computer-aided diagnosis. Scientific reports, 12(1), 1762. https://www.nature.com/articles/s41598-022-05725-7
15. Keles, E., & Bagci, U. (2023). The past, current, and future of neonatal intensive care units with artificial intelligence: a systematic review. NPJ digital medicine, 6(1), 220. https://www.nature.com/articles/s41746-023-00941-5
16. Day, L. T., Gore-Langton, G. R., Rahman, A. E., Basnet, O., Shabani, J., Tahsina, T., ... & Lawn, J. E. (2020). Labour and delivery ward register data availability, quality, and utility-Every Newborn-birth indicators research tracking in hospitals (EN-BIRTH) study baseline analysis in three countries. BMC health services research, 20(1), 737. https://link.springer.com/article/10.1186/s12913-020-5028-7
17. Bhanot, K. (2023). Synthetic data generation and evaluation for fairness (Doctoral dissertation, Rensselaer Polytechnic Institute). https://search.proquest.com/openview/56fc52700b150cc4920289ab6be0865b/1?pq-origsite=gscholar&cbl=18750&diss=y
18. Chen, R. J., Wang, J. J., Williamson, D. F., Chen, T. Y., Lipkova, J., Lu, M. Y., ... & Mahmood, F. (2023). Algorithmic fairness in artificial intelligence for medicine and healthcare. Nature biomedical engineering, 7(6), 719-742. https://www.nature.com/articles/s41551-023-01056-8
19. Sacco, S. J., Chen, K., Wang, F., & Aseltine, R. (2023). Target-based fusion using social determinants of health to enhance suicide prediction with electronic health records. PloS one, 18(4), e0283595. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0283595.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
You are free to:
- Share — copy and redistribute the material in any medium or format
- Adapt — remix, transform, and build upon the material for any purpose, even commercially.
Terms:
- Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits.