Author:
Mikhaylov D V,Emelyanov G M
Abstract
Abstract
The offered paper is devoted to the problem of oneness and integrity of image for the semantic pattern (i.e., sense standard) revealed phrase by phrase for some text within a topical collection. One phrase corresponds here to an extended natural-language sentence. The basis of estimating affinity to the standard is the classifying of words of each phrase in a text according to the TF-IDF value relative to some text corpus. Texts to the corpus are pre-selected by an expert. The essence of the problem: for each phrase, its maximal affinity to the sense standard is achieved concerning the individual corpus document, and, consequently, it is necessary to estimate the mutual relevance of such documents concerning different phrases of the analyzed text. Based on distances between vectors of TF-IDF for words of a separate phrase obtained relative to different corpus documents, the significance estimation for each such document is entered into consideration to choose a pair of mutual relevant.
Subject
General Physics and Astronomy
Reference8 articles.
1. Estimation of the closeness to a semantic pattern of a topical text without construction of periphrases;Mikhaylov;Pattern Recognition and Image Analysis,2019
2. A statistical interpretation of term specificity and its application in retrieval;Jones;Journal of Documentation,2004
3. Hierarchization of topical texts based on the estimate of proximity to the semantic pattern without paraphrasing;Mikhaylov;Pattern Recognition and Image Analysis,2020
4. Formation of the representation of topical knowledge units in the problem of their estimation on the basis of open tests;Emelyanov;Machine learning and data analysis,2014