Named-Entity Recognition for Hindi language using context pattern-based maximum entropy

Author:

Jain Arti,Yadav Divakar,Tayal Devendra Kr,Arora Anuja

Abstract

This paper describes Named Entity Recognition (NER) system for Hindi language using two methodologies. An existing BaseLine Maximum Entropy-based Named Entity (BL-MENE) model and Context Pattern-based MENE (CP-MENE) framework the one proposed in this work. BL-MENE utilizes several features for the NER task but suffers from inaccurate Named Entity (NE) boundary detection, mis-classification errors, and partial recognition of NEs due to certain missing essentials. However, CP-MENE based NER task incorporates extensive features and patterns set to overcome these problems. In fact, the CP-MENE features include right-boundary, left-boundary, part-of-speech, synonyms, gazetteers and relative pronoun features. CP-MENE formulates a kind of recursive relationship to extract high ranked NE patterns that are generated through regular expressions via python@ code. Nowadays, since the Web contents in the Hindi language are rising, especially in the health-care applications, this work is conducted on the Hindi Health Data (HHD) corpus at Kaggle dataset. We conducted experiments on four NE categories- Person (PER), Disease (DIS), Consumable (CNS) and Symptom (SMP). Usually, researchers’ work upon PER NE within news articles while other NEs, especially related to the health-care domain such as DIS, CNS, and SMP NE types are left out which are incorporated in this research. CP-MENE improvised the classification performance of NEs and the F-measure achieved are 79.68% for PER, 72.50% for DIS, 68.78% for CNS, and 67.23% for SMP respectively which are comparable with respect to other NER approaches.

Publisher

AGHU University of Science and Technology Press

Subject

Artificial Intelligence,Computational Theory and Mathematics,Computer Graphics and Computer-Aided Design,Computer Networks and Communications,Computer Vision and Pattern Recognition,Modeling and Simulation,Computer Science (miscellaneous)

Cited by 5 articles. 订阅此论文施引文献 订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3