Virus de ácido ribonucleico (ARN) y coronavirus en Google Dataset Search: alcance y correlación epidemiológica-Reference-Cited by-同舟云学术

Virus de ácido ribonucleico (ARN) y coronavirus en Google Dataset Search: alcance y correlación epidemiológica

Published:2020-12-21 Issue: Volume: Page:
ISSN:1386-6710
Container-title:El profesional de la información
language:
Short-container-title:EPI

Author:

Blázquez-Ochando Manuel¹^ORCID,Prieto-Gutiérrez Juan-José¹^ORCID

Affiliation:

1. Universidad Complutense de Madrid

Abstract

This paper presents an analysis of the publication of datasets collected via Google Dataset Search, specialized in families of RNA viruses, whose terminology was obtained from the National Cancer Institute (NCI) thesaurus developed by the US Department of Health and Human Services. The objective is to determine the scope and reuse capacity of the available data, determine the number of datasets and their free access, the proportion in reusable download formats, the main providers, their publication chronology, and to verify their scientific provenance. On the other hand, we also define possible relationships between the publication of datasets and the main pandemics that have occurred during the last 10 years. The results obtained highlight that only 52% of the datasets are related to scientific research, while an even smaller fraction (15%) are reusable. There is also an upward trend in the publication of datasets, especially related to the impact of the main epidemics, as clearly confirmed for the Ebola virus, Zika, SARS-CoV, H1N1, H1N5, and especially the SARS-CoV-2 coronavirus. Finally, it is observed that the search engine has not yet implemented adequate methods for filtering and monitoring the datasets. These results reveal some of the difficulties facing open science in the dataset field. Resumen Se presenta un análisis sobre la publicación de conjuntos de datos recogidos en el buscador Google Dataset Search, especializados en familias de virus de ARN, cuya terminología fue obtenida en el tesauro del National Cancer Institute (NCI), elaborado por el Department of Health and Human Services de los Estados Unidos. Se busca evaluar el alcance y capacidad de reutilización de los datos disponibles, determinando el número de datasets, su libre acceso, proporción en formatos de descarga reutilizables, principales proveedores, cronología de publicación y verificación de su procedencia científica. Por otra parte, definir posibles vínculos entre la publicación de datasets y las principales pandemias ocurridas en los últimos 10 años. Entre los resultados obtenidos se destaca que sólo el 52% de los datasets tienen correspondencia con investigaciones científicas y, en menor medida, un 15% son reaprovechables. También se observa una evolución al alza en la publicación de datasets, especialmente vinculada a la afectación de las principales epidemias. Esto es confirmado de manera evidente con los virus del Ébola, Zika, SARS-CoV, H1N1, H1N5 y, particularmente con el coronavirus SARS-CoV-2. Finalmente, se observa que el buscador aún no ha implementado métodos adecuados para el filtrado y supervisión de los datasets. Estos resultados muestran algunas de las dificultades que aún presenta la ciencia abierta en el campo de los datasets.

Publisher

Ediciones Profesionales de la Informacion SL

Subject

Library and Information Sciences,Information Systems

Reference40 articles.

1. Ahlawat, Khyati; Chug, Anuradha; Singh, Amit-Prakash (2019). “Empirical evaluation of Map Reduce based hybrid approach for problem of imbalanced classification in big data”. International journal of grid and high performance computing, v. 11, n. 3, pp. 23-45. https://doi.org/10.4018/IJGHPC.2019070102

2. Bekelman, Justin E.; MPhil, Yan-Li; Gross, Cary P. (2003). “Scope and impact of financial conflicts of interest in biomedical research: a systematic review”. Jama, v. 289, n. 4, pp. 454-465. https://doi.org/10.1001/jama.289.4.454

3. Blischak, John D.; Davenport, Emily R.; Wilson, Greg (2016). “A quick introduction to version control with Git and GitHub”. PLoS computational biology, v. 12, n. 1. https://doi.org/10.1371/journal.pcbi.1004668

4. Brickley, Dan; Burgess, Matthew; Noy, Natasha (2019). “Google Dataset Search: Building a search engine for datasets in an open web ecosystem”. In: Proceedings of the 19th World wide web conference (WWW’19), pp. 1365-1375. https://doi.org/10.1145/3308558.3313685

5. Broder, Andrei (2002). “A taxonomy of web search”. ACM Sigir forum, v. 36, n. 2, pp. 3-10. https://doi.org/10.1145/792550.792552