What do pre-trained code models know about code? (ASE 2021 - New Ideas and Emerging Results (NIER) track)

Write a Blog >>

Sun 14 - Sat 20 November 2021 Australia

Who

Anjan Karmakar, Romain Robbes

Track

ASE 2021 NIER track

Time Zone

The program is currently displayed in (GMT+11:00) Hobart.

Use conference time zone: (GMT+11:00) HobartSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Wed 17 Nov 2021 09:40 - 09:50 at Kangaroo - Learning I Chair(s): Denys Poshyvanyk

Abstract

Pre-trained models of code built on the transformer architecture have performed well on many software engineering (SE) tasks, including predictive code generation. However, whether the vector representations from these pre-trained models comprehensively encode characteristics of source code well enough to be applicable to a broad spectrum of downstream tasks remains an open question.

One way to investigate this is with diagnostic tasks called probes. In this paper, we construct four probing tasks (probing for surface-level, syntactic, structural, and semantic information) for pre-trained code models. We show how probes can be used to identify whether models are deficient in (understanding) certain code properties, characterize different model layers, and get insight into the model sample-efficiency that may be necessary for each type of task.

We probe four models that vary in their expected knowledge of code properties: BERT (pre-trained on English), CodeBERT and CodeBERTa (pre-trained on source code, and natural language documentation), and GraphCodeBERT (pre-trained on source code with dataflow). While GraphCodeBERT performs more consistently overall, we find that BERT performs surprisingly well on some code tasks, which calls for further investigation. We release all the task datasets and evaluation code publicly.

Anjan Karmakar

Free University of Bozen-Bolzano

Romain Robbes