Semi-automatic Extraction of Plants Morphological Characters from Taxonomic Descriptions Written in Spanish-Reference-Cited by-同舟云学术

Semi-automatic Extraction of Plants Morphological Characters from Taxonomic Descriptions Written in Spanish

Published:2018-06-26 Issue: Volume:6 Page:e21282
ISSN:1314-2828
Container-title:Biodiversity Data Journal
language:
Short-container-title:BDJ

Author:

Mora Maria,Araya José

Abstract

Taxonomic literature keeps records of the planet's biodiversity and gives access to the knowledge needed for its sustainable management. Unfortunately, most of the taxonomic information is available in scientific publications in text format. The amount of publications generated is very large; therefore, to process it in order to obtain high structured texts would be complex and very expensive. Approaches like citizen science may help the process by selecting whole fragments of texts dealing with morphological descriptions; but a deeper analysis, compatible with accepted ontologies, will require specialised tools. The Biodiversity Heritage Library (BHL) estimates that there are more than 120 million pages published in over 5.4 million books since 1469, plus about 800,000 monographs and 40,000 journal titles (12,500 of these are current titles).It is necessary to develop standards and software tools to extract, integrate and publish this information into existing free and open access repositories of biodiversity knowledge to support science, education and biodiversity conservation.This document presents an algorithm based on computational linguistics techniques to extract structured information from morphological descriptions of plants written in Spanish. The developed algorithm is based on the work of Dr. Hong Cui from the University of Arizona; it uses semantic analysis, ontologies and a repository of knowledge acquired from the same descriptions. The algorithm was applied to the books Trees of Costa Rica Volume III (TCRv3), Trees of Costa Rica Volume IV (TCRv4) and to a subset of descriptions of the Manual of Plants of Costa Rica (MPCR) with very competitive results (more than 92.5% of average performance). The system receives the morphological descriptions in tabular format and generates XML documents. The XML schema allows documenting structures, characters and relations between characters and structures. Each extracted object is associated with attributes like name, value, modifiers, restrictions, ontology term id, amongst other attributes.The implemented tool is free software. It was developed using Java and integrates existing technology as FreeLing, the Plant Ontology (PO), the Plant Glossary, the Ontology Term Organizer (OTO) and the Flora Mesoamericana English-Spanish Glossary.

Publisher

Pensoft Publishers

Subject

Ecology,Ecology, Evolution, Behavior and Systematics

Link

https://bdj.pensoft.net/lib/ajax_srv/article_elements_srv.php?action=download_pdf&item_id=21282

Reference24 articles.

1. The Plant Ontology Database: a community resource for plant structure and developmental stages controlled vocabulary and annotations

2. Numbers of Living Species in Australia and the World;Chapman;Heritage,2009

3. CharaParser for fine-grained semantic annotation of organism morphological descriptions

4. Building the "Plant Glossary"—A controlled botanical vocabulary using terms extracted from the Floras of North America and China

5. Glosario Inglés-Español, Español-Inglés para Flora Mesoamericana.;Fernando

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Using natural language processing to extract plant functional traits from unstructured text;2023-11-06

2. Structuring Information from Plant Morphological Descriptions using Open Information Extraction;Biodiversity Information Science and Standards;2023-09-21

3. Fungal Dispersal Across Spatial Scales;Annual Review of Ecology, Evolution, and Systematics;2022-11-02

4. Essential Biodiversity Variables: extracting plant phenological data from specimen labels using machine learning;Research Ideas and Outcomes;2022-08-23