COMCAT: Leveraging Human Judgment to Improve Automatic Documentation and   Summarization

Skyler Grandel (1); Scott Thomas Andersen (2); Yu Huang (1); Kevin; Leach (1) ((1) Vanderbilt University; (2) Universidad Nacional Aut\`onoma de; M\`exico)

arXiv:2407.13648·cs.SE·July 19, 2024·1 cites

COMCAT: Leveraging Human Judgment to Improve Automatic Documentation and Summarization

Skyler Grandel (1), Scott Thomas Andersen (2), Yu Huang (1), Kevin, Leach (1) ((1) Vanderbilt University, (2) Universidad Nacional Aut\`onoma de, M\`exico)

PDF

Open Access

TL;DR

COMCAT enhances code comprehension by automatically generating relevant comments using augmented large language models, significantly outperforming standard methods and matching human quality in a variety of software engineering tasks.

Contribution

This paper introduces COMCAT, a novel LLM-based approach that intelligently generates and selects comments to improve software understanding, with a new dataset for further research.

Findings

01

COMCAT-generated comments significantly improve developer comprehension by up to 12%.

02

Participants preferred COMCAT comments over standard ChatGPT comments in 92% of cases.

03

COMCAT comments are as accurate and readable as human comments.

Abstract

Software maintenance constitutes a substantial portion of the total lifetime costs of software, with a significant portion attributed to code comprehension. Software comprehension is eased by documentation such as comments that summarize and explain code. We present COMCAT, an approach to automate comment generation by augmenting Large Language Models (LLMs) with expertise-guided context to target the annotation of source code with comments that improve comprehension. Our approach enables the selection of the most relevant and informative comments for a given snippet or file containing source code. We develop the COMCAT pipeline to comment C/C++ files by (1) automatically identifying suitable locations in which to place comments, (2) predicting the most helpful type of comment for each location, and (3) generating a comment based on the selected location and comment type. In a human…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Data Quality and Management · Natural Language Processing Techniques