Styler: learning formatting conventions to repair Checkstyle violations-Reference-Cited by-同舟云学术

Styler: learning formatting conventions to repair Checkstyle violations

Published:2022-08-06 Issue:6 Volume:27 Page:
ISSN:1382-3256
Container-title:Empirical Software Engineering
language:en
Short-container-title:Empir Software Eng

Author:

Loriot Benjamin,Madeiral Fernanda^ORCID,Monperrus Martin

Abstract

AbstractEnsuring the consistent usage of formatting conventions is an important aspect of modern software quality assurance. To do so, the source code of a project should be checked against the formatting conventions (or rules) adopted by its development team, and then the detected violations should be repaired if any. While the former task can be automatically done by format checkers implemented in linters, there is no satisfactory solution for the latter. Manually fixing formatting convention violations is a waste of developer time and code formatters do not take into account the conventions adopted and configured by developers for the used linter. In this paper, we present Styler, a tool dedicated to fixing formatting rule violations raised by format checkers using a machine learning approach. For a given project, Styler first generates training data by injecting violations of the project-specific rules in violation-free source code files. Then, it learns fixes by feeding long short-term memory neural networks with the training data encoded into token sequences. Finally, it predicts fixes for real formatting violations with the trained models. Currently, Styler supports a single checker, Checkstyle, which is a highly configurable and popular format checker for Java. In an empirical evaluation, Styler repaired 41% of 26,791 Checkstyle violations mined from 104 GitHub projects. Moreover, we compared Styler with the IntelliJ plugin CheckStyle-IDEA and the machine-learning-based code formatters Naturalize and CodeBuff. We found out that Styler fixes violations of a diverse set of Checkstyle rules (24/25 rules), generates smaller repairs in comparison to the other systems, and predicts repairs in seconds once trained on a project. Through a manual analysis, we identified cases in which Styler does not succeed to generate correct repairs, which can guide further improvements in Styler. Finally, the results suggest that Styler can be useful to help developers repair Checkstyle formatting violations.

Funder

Royal Institute of Technology

Publisher

Springer Science and Business Media LLC

Subject

Software

Link

https://link.springer.com/content/pdf/10.1007/s10664-021-10107-0.pdf

Reference30 articles.

1. Aftandilian E, Sauciuc R, Priya S, Krishnan S (2012) Building Useful Program Analysis Tools Using an Extensible Java Compiler. In: Proceedings of the 12th IEEE International Working Conference on Source Code Analysis and Manipulation (SCAM ’12), IEEE Computer Society, USA, pp 14–23 https://doi.org/10.1109/SCAM.2012.28

2. Ahmed UZ, Kumar P, Karkare A, Kar P, Gulwani S (2018) Compilation Error Repair: For the Student Programs, From the Student Programs. In: Proceedings of the 40th International Conference on Software Engineering: Software Engineering Education and Training (ICSE-SEET ’18), Association for Computing Machinery, New York, NY, USA, pp 78–87, https://doi.org/10.1145/3183377.3183383

3. Allamanis M, Barr ET, Bird C, Sutton C (2014) Learning Natural Coding Conventions. In: Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE ’14), Association for Computing Machinery, New York, NY, USA, pp 281–293 https://doi.org/10.1145/2635868.2635883https://doi.org/10.1145/2635868.2635883

4. Ayewah N, Hovemeyer D, Morgenthaler JD, Penix J, Pugh W (2008) Using Static Analysis to Find Bugs. IEEE Software 25(5):22–29. https://doi.org/10.1109/MS.2008.130

5. Bader J, Scott A, Pradel M, Chandra S (2019) Getafix: Learning to Fix Bugs Automatically. Proceedings of the ACM on Programming Languages 3(OOPSLA) https://doi.org/10.1145/3360585

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Enhancing Code Readability through Automated Consistent Formatting;Electronics;2024-05-27

2. From Leaks to Fixes: Automated Repairs for Resource Leak Warnings;Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering;2023-11-30

3. Using Deep Learning to Automatically Improve Code Readability;2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE);2023-09-11

4. MLinter: Learning Coding Practices from Examples—Dream or Reality?;2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER);2023-03

5. On the diffusion of test smells and their relationship with test code quality of Java projects;Journal of Software: Evolution and Process;2023-01-18