Abstract
Background: Electronic health records (EHRs) contain vast amounts of unstructured clinical narrative data that remain underutilized for pediatric research and quality improvement. Objective: To develop and validate a natural language processing (NLP) pipeline for automated extraction of pediatric clinical phenotypes from free-text EHR data. Methods: A transformer-based NLP model (BioClinicalBERT fine-tuned on pediatric corpora) was developed using 18,432 annotated pediatric clinical notes. Performance was evaluated against manual chart review for 12 pediatric conditions. Results: The NLP pipeline achieved mean F1-score of 0.91 (range: 0.86–0.96) across all 12 conditions, with processing speed 847 times faster than manual review. Interrater agreement between NLP and expert reviewers was κ=0.89. Conclusion: Transformer-based NLP demonstrates high accuracy for automated pediatric clinical phenotyping, enabling scalable EHR-based research and quality improvement initiatives.
References

This work is licensed under a Creative Commons Attribution 4.0 International License.
