The epidemiological characteristics of stroke phenotypes defined with ICD-10 and free-text: a cohort study linked to electronic health records-Reference-Cited by-同舟云学术

The epidemiological characteristics of stroke phenotypes defined with ICD-10 and free-text: a cohort study linked to electronic health records

Published:2023-04-04 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Davidson Emma M^ORCID,Casey Arlene,Grover Claire,Alex Beatrice,Wu Honghan,Campbell Archie^ORCID,Chalmers Fionna,Adams Mark,Iveson Matthew^ORCID,McIntosh Andrew M^ORCID,Ball Emily^ORCID,Rannikmae Kristiina^ORCID,Whalley Heather^ORCID,Whiteley William N^ORCID

Abstract

AbstractBackgroundCoded healthcare data may not capture all stroke cases and has limited accuracy for stroke subtypes. We sought to determine the incremental value of adding natural language processing (NLP) of free-text radiology reports to international classification of disease (ICD-10) codes to phenotype stroke, and stroke subtypes, in routinely collected healthcare datasets.MethodsWe linked participants in a community-based prospective cohort study, Generation Scotland, to clinical brain imaging reports (2008-2020) from five Scottish health boards. We used five combinations of NLP outputs and ICD-10 codes to define stroke phenotypes. With these phenotype models we measured the: stroke incidence standardised to a European Standardised Population; adjusted hazard ratio (aHR) of baseline hypertension for later stroke; and proportion of participants allocated stroke subtypes.ResultsOf 19,026 participants, over a mean follow-up of 10.2 years, 1938 had 3493 brain scans. Any stroke was identified in 534 participants: 319 with NLP alone, 59 with ICD-10 codes alone and 156 with both ICD-10 codes and an NLP report consistent with stroke. The stroke aHR for baseline hypertension was 1.47 (95%CI: 1.12-1.92) for NLP-defined stroke only; 1.57 (95%CI: 1.18-2.10) for ICD-10 defined stroke only; and 1.81 (95%CI: 1.20-2.72) for cases with ICD 10 stroke codes and NLP stroke phenotypes. The age-standardised incidence of stroke for these phenotype models was 1.35, 1.34, and 0.65 per 1000 person years, respectively. The proportion of strokes not subtyped was 26% (57/215) using only ICD-10, 9% (42/467) using only NLP, and 12% (65/534) using both NLP and ICD-10.ConclusionsAddition of NLP derived phenotypes to ICD-10 stroke codes identified approximately 2.5 times more stroke cases and greatly increased the proportion with subtyping. The phenotype model using ICD 10 stroke codes and NLP stroke phenotypes had the strongest association with baseline hypertension. This information is relevant to large cohort studies and clinical trials that use routine electronic health records for outcome ascertainment.

Publisher

Cold Spring Harbor Laboratory

Reference25 articles.

1. Accuracy of identifying incident stroke cases from linked health care data in UK Biobank

2. Text mining brain imaging reports

3. SemEHR: A general-purpose semantic search system to surface semantic data from clinical notes for tailored care, trial recruitment, and clinical research*;Journal of the American Medical Informatics Association,2018

4. S Fu , Leung L , Raulli A-O , Kallmes D , Kinsman K , Nelson K , et al. Assessment of the Impact of EHR Heterogeneity for Clinical Research Through A Case Study of Silent Brain Infarction. BMC Medical Informatics and Decision Making. 2020.

5. Developing automated methods for disease subtyping in UK Biobank: an exemplar study on stroke;BMC Medical Informatics and Decision Making,2021