Semantic key phrase-based model for document management-Reference-Cited by-同舟云学术

Semantic key phrase-based model for document management

Published:2019-06-19 Issue:6 Volume:26 Page:1709-1727
ISSN:1463-5771
Container-title:Benchmarking: An International Journal
language:en
Short-container-title:BIJ

Author:

Bafna Prafulla,Pramod Dhanya,Shrwaikar Shailaja,Hassan Atiya

Abstract

Purpose Document management is growing in importance proportionate to the growth of unstructured data, and its applications are increasing from process benchmarking to customer relationship management and so on. The purpose of this paper is to improve important components of document management that is keyword extraction and document clustering. It is achieved through knowledge extraction by updating the phrase document matrix. The objective is to manage documents by extending the phrase document matrix and achieve refined clusters. The study achieves consistency in cluster quality in spite of the increasing size of data set. Domain independence of the proposed method is tested and compared with other methods. Design/methodology/approach In this paper, a synset-based phrase document matrix construction method is proposed where semantically similar phrases are grouped to reduce the dimension curse. When a large collection of documents is to be processed, it includes some documents that are very much related to the topic of interest known as model documents and also the documents that deviate from the topic of interest. These non-relevant documents may affect the cluster quality. The first step in knowledge extraction from the unstructured textual data is converting it into structured form either as term frequency-inverse document frequency matrix or as phrase document matrix. Once in structured form, a range of mining algorithms from classification to clustering can be applied. Findings In the enhanced approach, the model documents are used to extract key phrases with synset groups, whereas the other documents participate in the construction of the feature matrix. It gives a better feature vector representation and improved cluster quality. Research limitations/implications Various applications that require managing of unstructured documents can use this approach by specifically incorporating the domain knowledge with a thesaurus. Practical implications Experiment pertaining to the academic domain is presented that categorizes research papers according to the context and topic, and this will help academicians to organize and build knowledge in a better way. The grouping and feature extraction for resume data can facilitate the candidate selection process. Social implications Applications like knowledge management, clustering of search engine results, different recommender systems like hotel recommender, task recommender, and so on, will benefit from this study. Hence, the study contributes to improving document management in business domains or areas of interest of its users from various strata’s of society. Originality/value The study proposed an improvement to document management approach that can be applied in various domains. The efficacy of the proposed approach and its enhancement is validated on three different data sets of well-articulated documents from data sets such as biography, resume and research papers. These results can be used for benchmarking further work carried out in these areas.

Publisher

Emerald

Subject

Business and International Management,Strategy and Management

Reference59 articles.

1. Organization and technology in knowledge transfer;Benchmarking: An International Journal,2004

2. Semantic clustering driven approaches to recommender systems,2016

3. The peculiarities of the text document representation,2015

4. Keyphrase Extraction in scientific articles: a supervised approach,2012

5. Cambria, E., Poria, S., Bisio, F., Bajpai, R. and Chaturvedi, I. (2015), “The CLSA model: a novel framework for concept-level sentiment analysis”, Computational Linguistics and Intelligent Text Processing, Springer International Publishing, pp. 3-22.

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Implementation of a system for documentary procedures in public institutions applying Robotic Process Automation (RPA);2023 IEEE XXX International Conference on Electronics, Electrical Engineering and Computing (INTERCON);2023-11-02

2. XML CLUSTERING FRAMEWORK BASED ON DOCUMENT CONTENT AND STRUCTURE IN A HETEROGENEOUS DIGITAL LIBRARY;Malaysian Journal of Computer Science;2023-04-30