Are Human Rules Necessary? Generating Reusable APIs with CoT Reasoning and In-Context Learning-Reference-Cited by-同舟云学术

Are Human Rules Necessary? Generating Reusable APIs with CoT Reasoning and In-Context Learning

Published:2024-07-12 Issue:FSE Volume:1 Page:2355-2377
ISSN:2994-970X
Container-title:Proceedings of the ACM on Software Engineering
language:en
Short-container-title:Proc. ACM Softw. Eng.

Author:

Mai Yubo¹^ORCID,Gao Zhipeng²^ORCID,Hu Xing¹^ORCID,Bao Lingfeng¹^ORCID,Liu Yu¹^ORCID,Sun JianLing¹^ORCID

Affiliation:

1. Zhejiang University, Hangzhou, China

2. Shanghai Institute for Advanced Study of Zhejiang University, Shanghai, China

Abstract

Inspired by the great potential of Large Language Models (LLMs) for solving complex coding tasks, in this paper, we propose a novel approach, named Code2API, to automatically perform APIzation for Stack Overflow code snippets. Code2API does not require additional model training or any manual crafting rules and can be easily deployed on personal computers without relying on other external tools. Specifically, Code2API guides the LLMs through well-designed prompts to generate well-formed APIs for given code snippets. To elicit knowledge and logical reasoning from LLMs, we used chain-of-thought (CoT) reasoning and few-shot in-context learning, which can help the LLMs fully understand the APIzation task and solve it step by step in a manner similar to a developer. Our evaluations show that Code2API achieves a remarkable accuracy in identifying method parameters (65%) and return statements (66%) equivalent to human-generated ones, surpassing the current state-of-the-art approach, APIzator, by 15.0% and 16.5% respectively. Moreover, compared with APIzator, our user study demonstrates that Code2API exhibits superior performance in generating meaningful method names, even surpassing the human-level performance, and developers are more willing to use APIs generated by our approach, highlighting the applicability of our tool in practice. Finally, we successfully extend our framework to the Python dataset, achieving a comparable performance with Java, which verifies the generalizability of our tool.

Funder

Starry Night Science Fund of Zhejiang University Shanghai Institute for Advanced Study

National Key Research and Development Program of China

Shanghai Sailing Program

National Science Foundation of China

Fundamental Research Funds for the Central Universities

Ningbo Natural Science Foundation

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.1145/3660811

Reference69 articles.

1. Patrick Bareiß Beatriz Souza Marcelo d’Amorim and Michael Pradel. 2022. Code generation tools (almost) for free? a study of few-shot pre-trained language models on code. arXiv preprint arXiv:2206.01335.

2. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33 (2020), 1877–1901.

3. Kaibo Cao, Chunyang Chen, Sebastian Baltes, Christoph Treude, and Xiang Chen. 2021. Automated query reformulation for efficient search based on query logs from stack overflow. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). 1273–1285.

4. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, and Greg Brockman. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.

5. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, and Sebastian Gehrmann. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.