Automated Summarization of Stack Overflow Posts (ICSE 2023 - Technical Track)

Who

Bonan Kou, Muhao Chen, Tianyi Zhang

Track

ICSE 2023 Technical Track

Time Zone

The program is currently displayed in (GMT+10:00) Hobart.

Use conference time zone: (GMT+10:00) HobartSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Fri 19 May 2023 11:15 - 11:30 at Meeting Room 102 - Developers' forums Chair(s): Omar Haggag

Abstract

Software developers often resort to Stack Overflow (SO) to fill their programming needs. Given the abundance of relevant posts, navigating them and comparing different solutions is tedious and time-consuming. Recent work has proposed to automatically summarize SO posts to concise text to facilitate the navigation of SO posts. However, these techniques rely only on information retrieval methods or heuristics for text summarization, which is insufficient to handle the ambiguity and sophistication of natural language.

This paper presents a deep learning based framework called ASSORT for SO post summarization. ASSORT includes two complementary learning methods, ASSORT$S$ and ASSORT${IS}$, to address the lack of labeled training data for SO post summarization. ASSORT$S$ is designed to directly train a novel ensemble learning model with BERT embeddings and domain-specific features to account for the unique characteristics of SO posts. By contrast, ASSORT${IS}$ is designed to reuse pre-trained models while addressing the domain shift challenge when no training data is present (i.e., zero-shot learning). Both ASSORT$S$ and ASSORT${IS}$ outperform six existing techniques by at least 13% and 7% respectively in terms of the F1 score. Furthermore, a human study shows that participants significantly preferred summaries generated by ASSORT$S$ and ASSORT${IS}$ over the best baseline, while the preference difference between ASSORT$S$ and ASSORT${IS}$ was small.

Bonan Kou

Purdue University

Muhao Chen

University of Southern California

Tianyi Zhang