Assessing and Restoring Reproducibility of Jupyter Notebooks (ASE 2020 - Research Papers)

Write a Blog >>

Mon 21 - Fri 25 September 2020 Melbourne, Australia

Who

Jiawei Wang, Tzu-yang Kuo, Li Li, Andreas Zeller

Track

ASE 2020 Research Papers

Time Zone

The program is currently displayed in (UTC) Coordinated Universal Time.

Use conference time zone: (UTC) Coordinated Universal TimeSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Tue 22 Sep 2020 08:40 - 09:00 at Kangaroo - Software Analysis (1) Chair(s): Michael Pradel

Abstract

Jupyter notebooks—documents that contain live code, equations, visualizations, and narrative text—now are among the most popular means to compute, present, discuss, and disseminate scientific findings. In principle, Jupyter notebooks should easily allow to reproduce and extend scientific computations and their findings; but in practice, this is not the case. The individual code cells in Jupyter notebooks can be executed in any order, with identifier usages preceding their definitions and results preceding their computations. In a sample of 936 published notebooks that would be executable in principle, we found that 73% of them would not be reproducible with straightforward approaches, requiring humans to infer (and often guess) the order in which the authors created the cells.

In this paper, we present an approach to

automatically satisfy dependencies between code cells to reconstruct possible execution orders of the cells; and
instrument code cells to mitigate the impact of non-reproducible statements (i.e., random functions) in Jupyter notebooks.

Our Osiris prototype takes a notebook as input and outputs the possible execution schemes that reproduce the exact notebook results. In our sample, Osiris was able to reconstruct such schemes for 82.23% of all executable notebooks, which has more than three times better than the state-of-the-art; the resulting reordered code is valid program code and thus available for further testing and analysis.

Jiawei Wang

Tzu-yang Kuo

The Hong Kong University of Science and Technology

Li Li

Monash University, Australia

Australia

Andreas Zeller

CISPA, Germany

Germany