GitRanking: A ranking of GitHub topics for software classification using active sampling-Reference-Cited by-同舟云学术

GitRanking: A ranking of GitHub topics for software classification using active sampling

Published:2023-07-18 Issue:10 Volume:53 Page:1982-2006
ISSN:0038-0644
Container-title:Software: Practice and Experience
language:en
Short-container-title:Softw Pract Exp

Author:

Sas Cezar¹^ORCID,Capiluppi Andrea¹^ORCID,Di Sipio Claudio²,Di Rocco Juri²,Di Ruscio Davide²

Affiliation:

1. Bernulli Institute University of Groningen Groningen The Netherlands

2. Department of Information Engineering Computer Science and Mathematics University of L'Aquila L'Aquila Italy

Abstract

AbstractContextGitHub is the world's most prominent host of source code, with more than 327M repositories. However, most of these repositories are not labelled or inadequately, making it harder for users to find relevant projects. Various proposals for software application domain classification over the past years have been proposed. However, these several of those approaches suffer from multiple issues, called antipatterns of software classification, that reduce their usability.ObjectiveIn this paper, we propose a new taxonomy in the GitHub ecosystem, called GitRanking, starting from a well‐structured data set, composed of curated repositories annotated with topics. The main objective is to create a baseline methodology for software classification that is expandable, hierarchical, grounded in a knowledge base, and free of antipatterns.MethodWe collected 121K topics from GitHub and used GitRanking to create a taxonomy of 301 ranked application domains. GitRanking (1) uses active sampling to ensure a minimal number of annotations to create the ranking; and (2) links each topic to Wikidata, reducing ambiguities and improving the reusability of the taxonomy. Furthermore, we adopt the conceived taxonomy in a classification task by considering a state‐of‐the‐art classifier.ResultsOur results show that GitRanking can effectively rank terms in a hierarchy according to how general or specific their meaning is. Furthermore, we show that GitRanking is a dynamically extensible method: it can currently accept further terms to be ranked, and with a minimum number of annotations (). Concerning the classification task, we show that the model achieves an F1‐score of 34%, with a precision of 54%.ConclusionThis paper is the first collective attempt at building a ground‐up taxonomy of software domains. Our vision is that our taxonomy, and its extensibility, can be used to better and more precisely label software projects.

Publisher

Wiley

Subject

Software

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/spe.3238

Reference48 articles.

1. SharmaA ThungF KochharPS SulistyaA LoD.Cataloging Github repositories. Paper presented at: Proceedings of the 21st International Conference on Evaluation and Assessment in Software Engineering EASE'17. Association for Computing Machinery; 2017; New York NY:314‐319. doi:10.1145/3084226.3084287

2. Automated Tagging of Software Projects Using Bytecode and Dependencies (N)

3. HiGitClass: Keyword-Driven Hierarchical Classification of GitHub Repositories

4. Di SipioC RubeiR Di RuscioD NguyenPT.A multinomial naı̈ve bayesian (MNB) network to automatically recommend topics for github repositories. Paper presented at: Proceedings of the Evaluation and Assessment in Software Engineering EASE'20. Association for Computing Machinery; 2020; New York NY:71‐80. doi:10.1145/3383219.3383227

5. GHTRec: A Personalized Service to Recommend GitHub Trending Repositories for Developers

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Estimating Software Project Performance Using Factor Analysis and Sequential Equation Modelling;2024 5th International Conference on Image Processing and Capsule Networks (ICIPCN);2024-07-03

2. Automated categorization of pre-trained models in software engineering: A case study with a Hugging Face dataset;Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering;2024-06-18

3. Multi-granular software annotation using file-level weak labelling;Empirical Software Engineering;2023-11-30