On the Construction of Web NER Model Training Tool based on Distant Supervision-Reference-Cited by-同舟云学术

On the Construction of Web NER Model Training Tool based on Distant Supervision

Published:2020-11-30 Issue:6 Volume:19 Page:1-28
ISSN:2375-4699
Container-title:ACM Transactions on Asian and Low-Resource Language Information Processing
language:en
Short-container-title:ACM Trans. Asian Low-Resour. Lang. Inf. Process.

Author:

Chou Chien-Lung¹,Chang Chia-Hui¹^ORCID,Lin Yuan-Hao¹,Chien Kuo-Chun¹

Affiliation:

1. National Central University, Taiwan (R.O.C.)

Abstract

Named entity recognition (NER) is an important task in natural language understanding, as it extracts the key entities (person, organization, location, date, number, etc.) and objects (product, song, movie, activity name, etc.) mentioned in texts. However, existing natural language processing (NLP) tools (such as Stanford NER) recognize only general named entities or require annotated training examples and feature engineering for supervised model construction. Since not all languages or entities have public NER support, constructing a tool for NER model training is essential for low-resource language or entity information extraction. In this article, we study the problem of developing a tool to prepare training corpus from the Web with known seed entities for custom NER model training via distant supervision. The major challenge of automatic labeling lies in the long labeling time due to large corpus and seed entities as well as the concern to avoid false positive and false negative examples due to short and long seeds. To solve this problem, we adopt locality-sensitive hashing (LSH) for various length of seed entities. We conduct experiments on five types of entity recognition tasks, including Chinese person names, food names, locations, points of interest (POIs), and activity names to demonstrate the improvements with the proposed Web NER model construction tool. Because the training corpus is obtained by automatic labeling of the seed entity–related sentences, one could use either the entire corpus or the positive only sentences for model training. Based on the experimental results, we found the decision should depend on whether traditional linear chained conditional random fields (CRF) or deep neural network–based CRF is used for model training as well as the completeness of the provided seed list.

Funder

Ministry of Science and Technology, Taiwan

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3422817

Reference59 articles.

1. Automatic acquisition of named entity tagged corpus from world wide web

2. Freebase

3. RAPID detection of gene–gene interactions in genome-wide association studies

4. New Avenues in Opinion Mining and Sentiment Analysis

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Fine-Grained Meetup Events Extraction Through Context-Aware Event Argument Positioning and Recognition;2024-07-01

2. EventGo! Mining Events Through Semi-Supervised Event Title Recognition and Pattern-based Venue/Date Coupling;J INF SCI ENG;2023

3. Autonomous schema markups based on intelligent computing for search engine optimization;PeerJ Computer Science;2022-12-08