Joining metadata and textual features to advise administrative courts decisions: a cascading classifier approach-Reference-Cited by-同舟云学术

Joining metadata and textual features to advise administrative courts decisions: a cascading classifier approach

Published:2023-02-18 Issue: Volume: Page:
ISSN:0924-8463
Container-title:Artificial Intelligence and Law
language:en
Short-container-title:Artif Intell Law

Author:

Mentzingen Hugo^ORCID,Antonio Nuno^ORCID,Lobo Victor^ORCID

Abstract

AbstractDecisions of regulatory government bodies and courts affect many aspects of citizens’ lives. These organizations and courts are expected to provide timely and coherent decisions, although they struggle to keep up with the increasing demand. The ability of machine learning (ML) models to predict such decisions based on past cases under similar circumstances was assessed in some recent works. The dominant conclusion is that the prediction goal is achievable with high accuracy. Nevertheless, most of those works do not consider important aspects for ML models that can impact performance and affect real-world usefulness, such as consistency, out-of-sample applicability, generality, and explainability preservation. To our knowledge, none considered all those aspects, and no previous study addressed the joint use of metadata and text-extracted variables to predict administrative decisions. We propose a predictive model that addresses the abovementioned concerns based on a two-stage cascade classifier. The model employs a first-stage prediction based on textual features extracted from the original documents and a second-stage classifier that includes proceedings’ metadata. The study was conducted using time-based cross-validation, built on data available before the predicted judgment. It provides predictions as soon as the decision date is scheduled and only considers the first document in each proceeding, along with the metadata recorded when the infringement is first registered. Finally, the proposed model provides local explainability by preserving visibility on the textual features and employing the SHapley Additive exPlanations (SHAP). Our findings suggest that this cascade approach surpasses the standalone stages and achieves relatively high Precision and Recall when both text and metadata are available while preserving real-world usefulness. With a weighted F1 score of 0.900, the results outperform the text-only baseline by 1.24% and the metadata-only baseline by 5.63%, with better discriminative properties evaluated by the receiver operating characteristic and precision-recall curves.

Funder

Universidade Nova de Lisboa

Publisher

Springer Science and Business Media LLC

Subject

Law,Artificial Intelligence

Link

https://link.springer.com/content/pdf/10.1007/s10506-023-09348-9.pdf

Reference38 articles.

1. Aletras N, Tsarapatsanis D, Preoţiuc-Pietro D, Lampos V (2016) Predicting judicial decisions of the European court of human rights: a natural language processing perspective. PeerJ Comput Sci 2016(10):1–19. https://doi.org/10.7717/peerj-cs.93

2. Bibal A, Lognoul M, De Streel A, Frénay B (2021) Legal requirements on explainability in machine learning. Artif Intell Law 29(2):149–169. https://doi.org/10.1007/s10506-020-09270-4

3. Bird S, Klein E, Loper E (2009) Natural language processing with python. O’Reilly Med. https://doi.org/10.5555/1717171

4. Blei DM, Ng AY, Jordan MI (2003) Latent dirichlet allocation. J Mach Learn Res 3(4–5):993–1022. https://doi.org/10.1016/b978-0-12-411519-4.00006-9

5. Brill E (1992) A simple rule-based part of speech tagger. In: Proceedings of the third conference on applied natural language processing. Association for Computational Linguistics. https://doi.org/10.3115/974499.974526