Natural language inference for Malayalam language using language agnostic sentence representation-Reference-Cited by-同舟云学术

Natural language inference for Malayalam language using language agnostic sentence representation

Published:2021-05-04 Issue: Volume:7 Page:e508
ISSN:2376-5992
Container-title:PeerJ Computer Science
language:en
Short-container-title:

Author:

Renjit Sara,Idicula Sumam

Abstract

Natural language inference (NLI) is an essential subtask in many natural language processing applications. It is a directional relationship from premise to hypothesis. A pair of texts is defined as entailed if a text infers its meaning from the other text. The NLI is also known as textual entailment recognition, and it recognizes entailed and contradictory sentences in various NLP systems like Question Answering, Summarization and Information retrieval systems. This paper describes the NLI problem attempted for a low resource Indian language Malayalam, the regional language of Kerala. More than 30 million people speak this language. The paper is about the Malayalam NLI dataset, named MaNLI dataset, and its application of NLI in Malayalam language using different models, namely Doc2Vec (paragraph vector), fastText, BERT (Bidirectional Encoder Representation from Transformers), and LASER (Language Agnostic Sentence Representation). Our work attempts NLI in two ways, as binary classification and as multiclass classification. For both the classifications, LASER outperformed the other techniques. For multiclass classification, NLI using LASER based sentence embedding technique outperformed the other techniques by a significant margin of 12% accuracy. There was also an accuracy improvement of 9% for LASER based NLI system for binary classification over the other techniques.

Funder

Research fellowship from Kerala State Council for Science, Technology and Environment

Publisher

PeerJ

Subject

General Computer Science

Link

https://peerj.com/articles/cs-508.pdf

Reference40 articles.

1. ArbTE: Arabic textual entailment;Alabbas,2011

2. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond;Artetxe;Transactions of the Association for Computational Linguistics,2019

3. Enriching word vectors with subword information;Bojanowski;Transactions of the Association for Computational Linguistics,2017

4. Textual entailment at EVALITA 2009;Bos;Proceedings of EVALITA,2009

5. Alignment based approach for Arabic textual entailment;Boudaa;Procedia Computer Science,2019

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Textual Entailment Technique for the Bahasa Using BiLSTM;2022 International Seminar on Intelligent Technology and Its Applications (ISITIA);2022-07-20

2. Textual Inference Identification in the Malayalam Language Using Convolutional Neural Network;Lecture Notes in Electrical Engineering;2022