Affiliation:
1. School of Information Science and Engineering, Shandong Agricultural University, Tai’an 271018, China
2. Key Laboratory of Huang-Huai-Hai Smart Agricultural Technology of Ministry of Agriculture and Rural Affairs, Tai’an 271018, China
Abstract
Text matching promotes the research and application of deep understanding of text information, and it provides the basis for information retrieval, recommendation systems and natural language processing by exploring the similar structures in text data. Owning to the outstanding performance and automatically extract text features for the target, the methods based-pre-training models gradually become the mainstream. However, such models usually suffer from the disadvantages of slow retrieval speed and low running efficiency. On the other hand, previous text matching algorithms have mainly focused on horizontal domain research, and there are relatively few vertical domain algorithms for agricultural text, which need to be further investigated. To address this issue, a second-order text matching algorithm has been developed. This paper first obtains a large amount of text about typical agricultural crops and constructs a database by using web crawlers and querying relevant textbooks, etc. Then BM25 algorithm is used to generate a candidate set and BERT model is used to filter the optimal match based on the candidate set. Experiments have shown that the Precision@1 of this second-order algorithm can reach 88.34% on the dataset constructed in this paper, and the average time to match a piece of text is only 2.02 s. Compared with BERT model and BM25 algorithm, there is an increase of 8.81% and 13.73% in Precision@1 respectively. In terms of the average time required for matching a text, it is 55.2 s faster than BERT model and only 2 s slower than BM25 algorithm. It can improve the efficiency and accuracy of agricultural information retrieval, agricultural decision support, agricultural market analysis, etc., and promote the sustainable development of agriculture.
Funder
Shandong Provincial Natural Science Foundation
Open Project Foundation of Intelligent Information Processing Key Laboratory of Shanxi Province
Reference42 articles.
1. How important is agriculture in China’s economic growth?;Yao;Oxf. Dev. Stud.,2000
2. Lin, C.X., Ding, B., Han, J., Zhu, F., and Zhao, B. (2008, January 15–19). Text cube: Computing ir measures for multidimensional text database analysis. Proceedings of the 2008 8th IEEE International Conference on Data Mining, Pisa, Italy.
3. Press “a” for artificial intelligence in agriculture: A review;Awasthi;JOIV Int. J. Inform. Vis.,2020
4. An owl-based specification of database management systems;Buraga;Comput. Mater. Contin,2022
5. Wang, S., and Jiang, J. (2015). Learning natural language inference with LSTM. arXiv.