Lexicon Development for COVID-19-related Concepts Using Open-source Word Embedding Sources: An Intrinsic and Extrinsic Evaluation-Reference-Cited by-同舟云学术

Lexicon Development for COVID-19-related Concepts Using Open-source Word Embedding Sources: An Intrinsic and Extrinsic Evaluation

Published:2021-02-22 Issue:2 Volume:9 Page:e21679
ISSN:2291-9694
Container-title:JMIR Medical Informatics
language:en
Short-container-title:JMIR Med Inform

Author:

Parikh Soham^ORCID,Davoudi Anahita^ORCID,Yu Shun^ORCID,Giraldo Carolina^ORCID,Schriver Emily^ORCID,Mowery Danielle^ORCID

Abstract

Background Scientists are developing new computational methods and prediction models to better clinically understand COVID-19 prevalence, treatment efficacy, and patient outcomes. These efforts could be improved by leveraging documented COVID-19–related symptoms, findings, and disorders from clinical text sources in an electronic health record. Word embeddings can identify terms related to these clinical concepts from both the biomedical and nonbiomedical domains, and are being shared with the open-source community at large. However, it’s unclear how useful openly available word embeddings are for developing lexicons for COVID-19–related concepts. Objective Given an initial lexicon of COVID-19–related terms, this study aims to characterize the returned terms by similarity across various open-source word embeddings and determine common semantic and syntactic patterns between the COVID-19 queried terms and returned terms specific to the word embedding source. Methods We compared seven openly available word embedding sources. Using a series of COVID-19–related terms for associated symptoms, findings, and disorders, we conducted an interannotator agreement study to determine how accurately the most similar returned terms could be classified according to semantic types by three annotators. We conducted a qualitative study of COVID-19 queried terms and their returned terms to detect informative patterns for constructing lexicons. We demonstrated the utility of applying such learned synonyms to discharge summaries by reporting the proportion of patients identified by concept among three patient cohorts: pneumonia (n=6410), acute respiratory distress syndrome (n=8647), and COVID-19 (n=2397). Results We observed high pairwise interannotator agreement (Cohen kappa) for symptoms (0.86-0.99), findings (0.93-0.99), and disorders (0.93-0.99). Word embedding sources generated based on characters tend to return more synonyms (mean count of 7.2 synonyms) compared to token-based embedding sources (mean counts range from 2.0 to 3.4). Word embedding sources queried using a qualifier term (eg, dry cough or muscle pain) more often returned qualifiers of the similar semantic type (eg, “dry” returns consistency qualifiers like “wet” and “runny”) compared to a single term (eg, cough or pain) queries. A higher proportion of patients had documented fever (0.61-0.84), cough (0.41-0.55), shortness of breath (0.40-0.59), and hypoxia (0.51-0.56) retrieved than other clinical features. Terms for dry cough returned a higher proportion of patients with COVID-19 (0.07) than the pneumonia (0.05) and acute respiratory distress syndrome (0.03) populations. Conclusions Word embeddings are valuable technology for learning related terms, including synonyms. When leveraging openly available word embedding sources, choices made for the construction of the word embeddings can significantly influence the words learned.

Publisher

JMIR Publications Inc.

Subject

Health Information Management,Health Informatics

Reference48 articles.

1. Ideas for how informaticians can get involved with COVID-19 research

2. ChapmanAPetersonKTuranoABoxTWallaceKJonesMA natural language processing system for national COVID-19 surveillance in the US Department of Veterans Affairs20201st Workshop on NLP for COVID-19 at ACL 2020July 2020Online

3. Large-scale identification of patients with cerebral aneurysms using natural language processing

4. Effectiveness of Lexico-syntactic Pattern Matching for Ontology Enrichment with Clinical Documents

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Know an Emotion by the Company It Keeps: Word Embeddings from Reddit/Coronavirus;Applied Sciences;2023-05-31

2. Evidence for Telemedicine’s Ongoing Transformation of Health Care Delivery Since the Onset of COVID-19: Retrospective Observational Study;JMIR Formative Research;2022-10-14

3. Is telemedicine the new normal? Evidence for the ongoing transformation of healthcare delivery since the onset of COVID-19 (Preprint);2022-04-11

4. Is virtual care the new normal? Evidence supporting Covid-19’s durable transformation on healthcare delivery;2022-03-09