Named entity recognition in Bengali and Hindi using support vector machine-Reference-Cited by-同舟云学术

Named entity recognition in Bengali and Hindi using support vector machine

Published:2011-07-07 Issue:1 Volume:34 Page:35-67
ISSN:0378-4169
Container-title:Lingvisticæ Investigationes. International Journal of Linguistics and Language Resources
language:en
Short-container-title:LI

Author:

Ekbal Asif¹,Bandyopadhyay Sivaji

Affiliation:

1. Indian Institute of Technology Patna, Patna, India

Abstract

Named Entity Recognition (NER) aims to classify each word of a document into predefined target named entity (NE) classes and is nowadays considered to be fundamental for many Natural Language Processing (NLP) tasks such as information retrieval, machine translation, information extraction, question answering systems and others. This paper reports about the development of a NER system for Bengali and Hindi using Support Vector Machine (SVM). We have used the annotated corpora of 122,467 tokens of Bengali and 502,974 tokens of Hindi tagged with the twelve different NE classes, defined as part of the IJCNLP-08 NER Shared Task for South and South East Asian Languages (SSEAL). An appropriate tag conversion routine has been developed in order to convert the data into the forms tagged with the four NE tags, namely Person name, Location name, Organization name and Miscellaneous name. The system makes use of the different contextual information of the words along with the variety of orthographic word-level features that are helpful in predicting the different NE classes. The system has been tested with the gold standard test sets of 35K, and 38K tokens for Bengali, and Hindi, respectively. Evaluation results have demonstrated the overall recall, precision, and f-score values of 85.11%, 81.74%, and 83.39%, respectively, for Bengali and 82.76%, 77.81%, and 80.21%, respectively, for Hindi. Statistical analysis, ANOVA is performed to show that the improvement in the performance with the use of language dependent features is statistically significant over the language independent features for Bengali and Hindi both.

Publisher

John Benjamins Publishing Company

Subject

Linguistics and Language

Link

http://www.jbe-platform.com/deliver/fulltext/li.34.1.02ekb.pdf

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. DeepSpacy-NER: an efficient deep learning model for named entity recognition for Punjabi language;Evolving Systems;2022-08-03

2. Named-Entity Recognition for Hindi language using context pattern-based maximum entropy;Computer Science;2022-03-24

3. A Novel Conceptual Chatbot Architecture for the Sinhala Language – A Case Study on Food Ordering Scenario;2022 2nd International Conference on Advanced Research in Computing (ICARC);2022-02-23

4. Research Trends for Named Entity Recognition in Hindi Language;Data Visualization and Knowledge Engineering;2019-08-10

5. Named Entity System for Tweets in Hindi Language;International Journal of Intelligent Information Technologies;2018-10