DETEXA: declarative extensible text exploration and analysis through SQL-Reference-Cited by-同舟云学术

DETEXA: declarative extensible text exploration and analysis through SQL

Published:2023-05-10 Issue: Volume: Page:
ISSN:1432-5012
Container-title:International Journal on Digital Libraries
language:en
Short-container-title:Int J Digit Libr

Author:

Foufoulas Yannis^ORCID,Zacharia Eleni,Dimitropoulos Harry,Manola Natalia,Ioannidis Yannis

Abstract

AbstractMetadata enrichment through text mining techniques is becoming one of the most significant tasks in digital libraries. Due to the exponential increase of open access publications, several new challenges have emerged. Raw data are usually big, unstructured, and come from heterogeneous data sources. In this paper, we introduce a text analysis framework implemented in extended SQL that exploits the scalability characteristics of modern database management systems. The purpose of this framework is to provide the opportunity to build performant end-to-end text mining pipelines which include data harvesting, cleaning, processing, and text analysis at once. SQL is selected due to its declarative nature which offers fast experimentation and the ability to build APIs so that domain experts can edit text mining workflows via easy-to-use graphical interfaces. Our experimental analysis demonstrates that the proposed framework is very effective and achieves significant speedup, up to three times faster, in common use cases compared to other popular approaches.

Funder

Horizon 2020 Framework Programme

Publisher

Springer Science and Business Media LLC

Subject

Library and Information Sciences

Link

https://link.springer.com/content/pdf/10.1007/s00799-023-00358-1.pdf

Reference24 articles.

1. NLTK, https://www.nltk.org

2. PySpark, https://spark.apache.org/docs/latest/api/python/

3. Dask, https://dask.org

4. Raasveldt, M., Mühleisen, H.: Vectorized udfs in column-stores. In: Proceedings of the 28th International Conference on Scientific and Statistical Database Management (2016)

5. https://www.postgresql.org/docs/current/xfunc.html

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Methods for generation, recommendation, exploration and analysis of scholarly publications;International Journal on Digital Libraries;2024-09-03