Code reasoning has emerged as a new paradigm for evaluating LLMs’ performance on programming tasks. While existing techniques assess LLMs’ code reasoning abilities in an \emph{explicit} manner—typically through tasks such as input/output prediction—the extent to which LLMs can \emph{implicitly} integrate execution reasoning into tasks that require simulating the code execution remains largely unexplored. Moreover, the oversimplified nature of current evaluation strategies limits researchers’ ability to analyze the reasoning process of LLMs beyond their final predictions, and hinders a rigorous assessment of their generalizability to real-world projects.
In this research proposal, we emphasize the need for a comprehensive and realistic evaluation of LLMs’ code reasoning capabilities. To address the limitations of prior work, we outline three research directions: (1) challenging LLMs with code reasoning tasks at different levels, (2) assessing the quality of LLMs’ intermediate reasoning processes, and (3) evaluating LLMs’ code reasoning performance under real-world contexts.
Program Display Configuration
Tue 14 Apr
Displayed time zone: Brasilia, Distrito Federal, Brazilchange