ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Fri 17 Apr 2026 16:00 - 16:15 at Oceania IX - AI for Software Engineering 29 Chair(s): Tien N. Nguyen

We introduce Modelizer—a novel framework that, given a black-box program, learns a model from its input/output behavior using neural machine translation algorithms. The resulting model mocks the original program: Given an input, the model predicts the output that would have been produced by the program. However, the model is also reversible—that is, the model can predict the input that would have produced a given output. Finally, the model is differentiable and can be efficiently restricted to predict only a certain aspect of the program behavior. Modelizer uses grammars to synthesize and inputs and unsupervised tokenizers to decompose the resulting outputs, allowing it to learn sequence-to-sequence associations between token streams. Other than input grammars, Modelizer only requires the ability to execute the program. The resulting models are small, requiring fewer than 6.3 million parameters for languages such as Markdown or HTML; and they are accurate, achieving up to 95.4% accuracy and a BLEU score of 0.98 with standard error 0.04 in mocking real-world applications. As it learns from and predicts executions rather than code, Modelizer departs from the LLM-centric research trend, opening new opportunities for program-specific models that are fully tuned towards individual programs. Indeed, we foresee several applications of these models, especially as the output of the program can be any aspect of program behavior. Beyond mocking and predicting program behavior, the models can also synthesize inputs that are likely to produce a particular behavior, such as failures or coverage, thus assisting in program understanding and maintenance.

Fri 17 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

16:00 - 17:30
AI for Software Engineering 29Journal-first Papers / Research Track at Oceania IX
Chair(s): Tien N. Nguyen University of Texas at Dallas
16:00
15m
Talk
Learning Program Behavioral Models from Synthesized Input-Output Pairs
Journal-first Papers
Tural Mammadov CISPA Helmholtz Center for Information Security, Dietrich Klakow Saarland University, Alexander Koller Saarland University, Andreas Zeller CISPA Helmholtz Center for Information Security
16:15
15m
Talk
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
Journal-first Papers
Hassan Sartaj Simula Research Laboratory, Shaukat Ali Simula Research Laboratory and Oslo Metropolitan University, Julie Marie Gjøby Welfare Technologies Section, Oslo Kommune Helseetaten
16:30
15m
Talk
Change And Cover: Last-Mile, Pull Request-Based Regression Test Augmentation
Research Track
Zitong Zhou UCLA, Matteo Paltenghi University of Stuttgart, Miryung Kim UCLA and Amazon Web Services, Michael Pradel CISPA Helmholtz Center for Information Security
Link to publication Media Attached
16:45
15m
Talk
HarnessLLM: Rust Verification Harness Generation with Large Language Models
Research Track
Minghua Wang Ant Group, Yuwei Liu Ant Group, Lin Huang Ant Group
17:00
15m
Talk
Agentic Predicates Reasoning for Directed Fuzzing
Research Track
Jie Zhu University of Chicago, Chihao Shen University of Maryland, Ziyang Li Johns Hopkins University, Jiahao Yu Northwestern University, Yizheng Chen University of Maryland, Kexin Pei The University of Chicago
Pre-print
17:15
15m
Talk
Relax with Capybaras
Research Track

Media Attached