On the Privacy–Utility Trade-Off in Differentially Private Hierarchical Text Classification-Reference-Cited by-同舟云学术

On the Privacy–Utility Trade-Off in Differentially Private Hierarchical Text Classification

Published:2022-11-04 Issue:21 Volume:12 Page:11177
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Wunderlich Dominik,Bernau Daniel,Aldà Francesco,Parra-Arnau Javier,Strufe Thorsten

Abstract

Hierarchical text classification consists of classifying text documents into a hierarchy of classes and sub-classes. Although Artificial Neural Networks have proved useful to perform this task, unfortunately, they can leak training data information to adversaries due to training data memorization. Using differential privacy during model training can mitigate leakage attacks against trained models, enabling the models to be shared safely at the cost of reduced model accuracy. This work investigates the privacy–utility trade-off in hierarchical text classification with differential privacy guarantees, and it identifies neural network architectures that offer superior trade-offs. To this end, we use a white-box membership inference attack to empirically assess the information leakage of three widely used neural network architectures. We show that large differential privacy parameters already suffice to completely mitigate membership inference attacks, thus resulting only in a moderate decrease in model utility. More specifically, for large datasets with long texts, we observed Transformer-based models to achieve an overall favorable privacy–utility trade-off, while for smaller datasets with shorter texts, convolutional neural networks are preferable.

Funder

European Union

“la Caixa” Foundation

Alexander von Humboldt Post-Doctoral Fellowship

Spanish Government

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/12/21/11177/pdf

Reference55 articles.

1. Uncertainty in big data analytics: Survey, opportunities, and challenges;J. Big Data,2019

2. Taylor, C. (2022, April 06). What’s the Big Deal with Unstructured Data? 2013. Wired. Available online: https://www.wired.com/insights/2013/09/whats-the-big-deal-with-unstructured-data/.

3. Mao, Y., Tian, J., Han, J., and Ren, X. (2019, January 3–7). Hierarchical Text Classification with Reinforced Label Assignment. Proceedings of the Conference on Empirical Methods in Natural Language Processing, Hong Kong, China.

4. An evaluation of classification models for question topic categorization;J. Am. Soc. Inf. Sci. Technol.,2012

5. Agrawal, R., Gupta, A., Prabhu, Y., and Varma, M. (2013, January 13–17). Multi-label learning with millions of labels: Recommending advertiser bid phrases for web pages. Proceedings of the International Conference on World Wide Web, Rio de Janeiro, Brazil.

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A survey on membership inference attacks and defenses in machine learning;Journal of Information and Intelligence;2024-09

2. Towards privacy preserved document image classification: a comprehensive benchmark;International Journal on Document Analysis and Recognition (IJDAR);2024-06-18

3. Hierarchical Text Classification and Its Foundations: A Review of Current Research;Electronics;2024-03-25