Story Beyond the Eye: Glyph Positions Break PDF Text Redaction-Reference-Cited by-同舟云学术

Story Beyond the Eye: Glyph Positions Break PDF Text Redaction

Published:2023-07 Issue:3 Volume:2023 Page:43-61
ISSN:2299-0984
Container-title:Proceedings on Privacy Enhancing Technologies
language:
Short-container-title:PoPETs

Author:

Bland Maxwell¹,Iyer Anushya¹,Levchenko Kirill¹

Affiliation:

1. University of Illinois, Urbana-Champaign

Abstract

In this work we find that many current redactions of PDF text are insecure due to non-redacted character positioning information. In particular, subpixel-sized horizontal shifts in redacted and non-redacted characters can be recovered and used to effectively deredact first and last names. Unfortunately these findings affect redactions where the text underneath the black box is removed from the PDF. We demonstrate these findings by performing a comprehensive vulnerability assessment of common PDF redaction types. We examine 11 popular PDF redaction tools, including Adobe Acrobat, and find that they leak information about redacted text. We also effectively deredact hundreds of real-world PDF redactions, including those found in OIG investigation reports and FOIA responses. To correct the problem, we have released open source algorithms to fix vulnerable redactions and reduce the amount of information leaked by nonexcising redactions (where the text underneath the redaction is copy-pastable). We have also notified the developers of the studied redaction tools. We have notified the Office of Inspector General, the Free Law Project, PACER, Adobe, Microsoft, and the US Department of Justice. We are working with several of these groups to prevent our discoveries from being used for malicious purposes.

Publisher

Privacy Enhancing Technologies Symposium Advisory Board

Subject

General Medicine

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. RedactBuster: Entity Type Recognition from Redacted Documents;Lecture Notes in Computer Science;2024

2. Detection of Redacted Text in Legal Documents;Linking Theory and Practice of Digital Libraries;2023

3. Making PDFs Accessible for Visually Impaired Users (and Findable for Everybody Else);Linking Theory and Practice of Digital Libraries;2023