ESEIW 2026
Sun 4 - Fri 9 October 2026 München, Germany

This program is tentative and subject to change.

\textbf{Background:} Smart contract vulnerabilities pose significant security risks in blockchain systems. Automated static auditing tools are commonly used in practice, yet their effectiveness varies across vulnerability categories and at scale. Recently, large language model (LLM)-based tools have been proposed, but their performance is rarely evaluated against real-world contracts and alongside established analyzers. \textbf{Aim:} This paper presents a systematic benchmarking study of smart contract auditing tools under a single, explicitly defined task: identifying smart contract vulnerabilities in Solidity contracts. \textbf{Method:} We evaluate static and LLM-based tools ($n=7$) on two benchmarks: SmartBugs Curated ($n=143$) and a filtered FORGE subset of real-world audit-derived contracts ($n=173$). We use a tool-agnostic evaluation pipeline to normalize heterogeneous tool outputs, reporting accuracy, top-$K$ detection, execution robustness, and cross-dataset performance. \textbf{Results:} Our findings expose systematic trade-offs across auditing approaches. Static tools optimized for broad vulnerability coverage tend to over-report categories, potentially increasing triage burden, while exhibiting stability limitations on real-world contracts. LLM-driven auditors demonstrate strong coverage across many vulnerability categories, but similarly suffer from over-prediction. \textbf{Conclusions:} These results show that current smart contract auditing tools differ less in their ability to identify potential vulnerabilities than in their ability to do so precisely and reliably, motivating future auditing workflows for practical smart contract vulnerability detection.

This program is tentative and subject to change.

Fri 9 Oct

Displayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change

16:00 - 17:30
16:00
15m
Talk
Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency
ESEM - Technical Track
Junchi Liu Colorado State University, Ali Bigdeli Colorado State University, Roya Daneshi Colorado State University, Atu Ambala Colorado State University, Sudipto Ghosh Colorado State University, USA, Fabio Marcos De Abreu Santos Colorado State University, USA
16:15
15m
Talk
Do Smart Contract Auditing Results Transfer Across Datasets? A Two-Benchmark Empirical Study of Static and LLM-Based Security Tools
ESEM - Technical Track
Shawal Khalid Virginia Tech, Joaquin Tuckett Virginia Tech, Chris Brown Virginia Tech
16:30
15m
Talk
Beyond Heuristics? Rethinking Targeted Unit Test Generation in the Era of LLM Agents
ESEM - Technical Track
Jiayu Liu Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Yanjun Wu Institute of Software, Chinese Academy of Sciences
16:45
15m
Talk
Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models
ESEM - Technical Track
Nils Kiele University of Calgary, Zainab Saad University of Calgary, Zirui Wang University of Calgary, Steve Drew University of Calgary, Samira Ebrahimi Kahou University of Calgary
17:00
15m
Talk
Evaluating LLM-Based Test Generation for a Large Industrial C++ Database System. A Case Study on SAP HANA
ESEM - Software Engineering in Practice Track
Vekil Bekmyradov TH Köln, Gummersbach, Germany, Thomas Bach SAP, Alexander Berndt Heidelberg University, Heidelberg Germany, Noah C. Puetz TH Köln, Gummersbach, Germany, Bartosz Bogacz SAP, Walldorf, Germany, Thomas Bartz-Beielstein TH Köln, Gummersbach, Germany
17:15
10m
Talk
APCA in the Loop: An Empirical Study of In-loop Patch Correctness Assessment for APR
ESEM - Emerging Results, Vision, and Reflection Papers Track
Sahand Moslemi Yengejeh Bilkent University, CS Department, Anil Koyuncu Bilkent University