Speaking in Dialects: A Reusable Dataset of Real-World TOSCA Orchestration Topologies
This program is tentative and subject to change.
The Topology and Orchestration Specification for Cloud Applications (TOSCA) is the OASIS standard for portable cloud orchestration, yet a decade of adoption has produced a fragmented landscape of mutually incompatible dialects and, remarkably, no public dataset with which to study any of it. We address this with a curated dataset of 14,931 real-world TOSCA blueprints drawn from 260 open-source repositories, identified among 1,973 candidates through three complementary discovery strategies (repository keyword, file-content, and topic search). Because no single parser spans every dialect, we curate the corpus with a three-tier, parser-independent quality model, i.e. parseability, structural coherence, and version classification, and annotate each file with provenance, dialect and version labels, and structural-complexity metrics. The result is a multi-dialect snapshot of TOSCA as practitioners actually write it: TOSCA 2.0 dominates (54.1%), but Simple Profile 1.x (33.6%) and engine-specific DSLs such as Cloudify and Alien4Cloud endure, alongside a long tail of vendor and academic extensions missing from the official specifications. The dataset supports empirical studies of TOSCA evolution, cross-project reuse, structural quality, and version migration. Concretely, it can also serve as training and evaluation data for LLM-based blueprint generation, repair, and version 1.x to 2.0 migration, and as ground truth for dialect classification, clone detection, and cross-dialect parser benchmarking.
This program is tentative and subject to change.
Wed 16 SepDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
11:00 - 12:30 | Session 2 - AI Under the Microscope: Analytics, Quality & InsightsRegistered Reports / Research Papers Track / Visions and Emerging Results Track / Tool Demonstration and Data Showcase Track / Industry Track at A53S Theme: AI-Driven Software Development | ||
11:00 20mPaper | To Copilot and Beyond: 22 AI Systems Developers Want Built Industry Track Rudrajit Choudhuri Oregon State University, Christian Bird Microsoft Research, Carmen Badea Microsoft Research, Anita Sarma Oregon State University | ||
11:20 20mPaper | RCSAgent: An LLM Agent Approach to Reference-Counting Semantics Recovery Research Papers Track Rui Wang Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China, Xutong Ma State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Xuecheng Li Independent Researcher, Jian Zhang Institute of Software at Chinese Academy of Sciences; University of Chinese Academy of Sciences | ||
11:40 20mPaper | Fusing UI Structure and Semantics for Feature-oriented App Screen Retrieval and Clustering Research Papers Track Arun Krishna Vajjala George Mason University, Yanfu Yan William & Mary, Ajay Krishna Vajjala George Mason University, Shrunal Pothagoni George Mason University, Denys Poshyvanyk William & Mary, Kevin Moran University of Central Florida | ||
12:00 10mPaper | How Humans, Bots, and Agents Communicate About Vulnerabilities in Pull Requests Registered Reports Pien Rooijendijk Radboud University, Christoph Treude Singapore Management University, Mairieli Wessel Radboud University Pre-print | ||
12:10 10mShort-paper | Speaking in Dialects: A Reusable Dataset of Real-World TOSCA Orchestration Topologies Tool Demonstration and Data Showcase Track Marco Tonnarelli JADS - TU/e, Stefano Fossati JADS - TU/e, Indika Kumara Tilburg University, Damian Andrew Tamburri University of Sannio - JADS/NXP Semiconductors | ||
12:20 10mShort-paper | Hidden Amplifiers: Cross-Level Risk in Software Supply Chains Visions and Emerging Results Track Rakesh Podder Colorado State University, Rafael Fabian Gonzalez Arellano Colorado State University, Indrajit Ray Colorado State University Pre-print Media Attached File Attached | ||