MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format-Reference-Cited by-同舟云学术

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format

Published:2017-04-12 Issue: Volume:4 Page:1453
ISSN:2046-1402
Container-title:F1000Research
language:en
Short-container-title:F1000Res

Author:

Ahmed Zeeshan,Dandekar Thomas

Abstract

Published scientific literature contains millions of figures, including information about the results obtained from different scientific experiments e.g. PCR-ELISA data, microarray analysis, gel electrophoresis, mass spectrometry data, DNA/RNA sequencing, diagnostic imaging (CT/MRI and ultrasound scans), and medicinal imaging like electroencephalography (EEG), magnetoencephalography (MEG), echocardiography (ECG), positron-emission tomography (PET) images. The importance of biomedical figures has been widely recognized in scientific and medicine communities, as they play a vital role in providing major original data, experimental and computational results in concise form. One major challenge for implementing a system for scientific literature analysis is extracting and analyzing text and figures from published PDF files by physical and logical document analysis. Here we present a product line architecture based bioinformatics tool ‘Mining Scientific Literature (MSL)’, which supports the extraction of text and images by interpreting all kinds of published PDF files using advanced data mining and image processing techniques. It provides modules for the marginalization of extracted text based on different coordinates and keywords, visualization of extracted figures and extraction of embedded text from all kinds of biological and biomedical figures using applied Optimal Character Recognition (OCR). Moreover, for further analysis and usage, it generates the system’s output in different formats including text, PDF, XML and images files. Hence, MSL is an easy to install and use analysis tool to interpret published scientific literature in PDF format.

Publisher

F1000 Research Ltd

Subject

General Pharmacology, Toxicology and Pharmaceutics,General Immunology and Microbiology,General Biochemistry, Genetics and Molecular Biology,General Medicine

Link

https://f1000research.com/articles/4-1453/v2/pdf

Reference44 articles.

1. Biomedical language processing: what’s beyond PubMed?;L Hunter;Mol Cell.,2006

2. Xed: A New Tool for Extracting Hidden Structures from Electronic Documents;K Hadjar;International Workshop on Document Image Analysis for Libraries.,2004

3. Database resources of the National Center for Biotechnology Information.;E Sayers;Nucleic Acids Res.,2010

4. MiSearch adaptive pubMed search tool.;D States;Bioinformatics.,2009

5. MScanner: a classifier for retrieving Medline citations.;G Poulter;BMC Bioinformatics.,2008

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Mining biomedical images towards valuable information retrieval in biomedical and life sciences;Database;2016