Towards a Universal Code Formatter through Machine Learning (SLE 2016 - Research Papers)

Who

Terence Parr, Jurgen Vinju

Track

SLE 2016

Time Zone

The program is currently displayed in (GMT+01:00) Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna.

Use conference time zone: (GMT+01:00) Amsterdam, Berlin, Bern, Rome, Stockholm, ViennaSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Tue 1 Nov 2016 10:30 - 10:55 at Zürich 2 - Development Environments Chair(s): Anthony Sloane

Abstract

There are many declarative frameworks that allow us to implement code formatters relatively easily for any specific language, but constructing them is cumbersome. The first problem is that “everybody” wants to format their code differently, leading to either many formatter variants or a ridiculous number of configuration options. Second, the size of each implementation scales with a language’s grammar size, leading to hundreds of rules.

In this paper, we solve the formatter construction problem using a novel approach, one that automatically derives formatters for any given language without intervention from a language expert. We introduce a code formatter called CodeBuff that uses machine learning to abstract formatting rules from a representative corpus, using a carefully designed feature set. Our experiments on Java, SQL, and ANTLR grammars show that CodeBuff is efficient, has excellent accuracy, and is grammar invariant for a given language. It also generalizes to a 4th language tested during manuscript preparation.

Link to Preprint

https://arxiv.org/abs/1606.08866v1

DOI

https://doi.org/10.1145/2997364.2997383

File attachments

PDF of slides for Terence Parr's presentation at SLE16 (codebuff-slides.pdf)	1.28MiB

Terence Parr

University of San Francisco, USA

Jurgen Vinju

CWI, Netherlands

Towards a Universal Code Formatter through Machine Learning