Gender bias in legal corpora and debiasing it-Reference-Cited by-同舟云学术

Gender bias in legal corpora and debiasing it

Published:2022-03-30 Issue:2 Volume:29 Page:449-482
ISSN:1351-3249
Container-title:Natural Language Engineering
language:en
Short-container-title:Nat. Lang. Eng.

Author:

Sevim Nurullah,Şahinuç Furkan,Koç Aykut

Abstract

AbstractWord embeddings have become important building blocks that are used profoundly in natural language processing (NLP). Despite their several advantages, word embeddings can unintentionally accommodate some gender- and ethnicity-based biases that are present within the corpora they are trained on. Therefore, ethical concerns have been raised since word embeddings are extensively used in several high-level algorithms. Studying such biases and debiasing them have recently become an important research endeavor. Various studies have been conducted to measure the extent of bias that word embeddings capture and to eradicate them. Concurrently, as another subfield that has started to gain traction recently, the applications of NLP in the field of law have started to increase and develop rapidly. As law has a direct and utmost effect on people’s lives, the issues of bias for NLP applications in legal domain are certainly important. However, to the best of our knowledge, bias issues have not yet been studied in the context of legal corpora. In this article, we approach the gender bias problem from the scope of legal text processing domain. Word embedding models that are trained on corpora composed by legal documents and legislation from different countries have been utilized to measure and eliminate gender bias in legal documents. Several methods have been employed to reveal the degree of gender bias and observe its variations over countries. Moreover, a debiasing method has been used to neutralize unwanted bias. The preservation of semantic coherence of the debiased vector space has also been demonstrated by using high-level tasks. Finally, overall results and their implications have been discussed in the scope of NLP in legal domain.

Publisher

Cambridge University Press (CUP)

Subject

Artificial Intelligence,Linguistics and Language,Language and Linguistics,Software

Reference107 articles.

1. Deep Contextualized Word Representations

2. Using background knowledge in case-based legal reasoning: A computational model and an intelligent learning environment

3. Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change

4. Legal Area Classification: A Comparative Study of Text Classifiers on Singapore Supreme Court Judgments

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Gender Bias Detection in Court Decisions: A Brazilian Case Study;The 2024 ACM Conference on Fairness, Accountability, and Transparency;2024-06-03

2. Measuring and Mitigating Gender Bias in Legal Contextualized Language Models;ACM Transactions on Knowledge Discovery from Data;2023-10-18

3. Unveiling the Black Box: Investigating the Interplay between AI Technologies, Explainability, and Legal Implications;2023 8th International Conference on Computer Science and Engineering (UBMK);2023-09-13

4. Creating a Chinese gender lexicon for detecting gendered wording in job advertisements;Information Processing & Management;2023-09

5. A Transformer-Based Prior Legal Case Retrieval Method;2023 31st Signal Processing and Communications Applications Conference (SIU);2023-07-05