Beyond Text: Understanding How Multimodal Generative AI Impacts Students Learning Software Development (Extended Abstract)
Recent advances in multimodal generative AI (GenAI) extend student support beyond text-focused prompting to include audio dialogue, media upload, and real-time stream (\ i.e., live screen-sharing). Yet, the educational implications of these multimodal capabilities in software engineering (SE) development settings remain underexplored. This extended abstract presents a proposed study to investigate how students leverage multimodal GenAI while completing realistic SE tasks in an open-source software (OSS) workflow, and how multimodal interaction shapes task-solving processes and outcomes. We will conduct in-person think-aloud sessions (min. n = 25) in which participants use a multimodal GenAI tool (Google AI Studio) across interaction modes while working through bug fixing, feature implementation with pull request contribution, and test writing tasks. We will triangulate screen/audio recordings, interaction traces, and pre/post surveys to characterize modality selection and prompting strategies, assess participants’ usage patterns, and explore impacts on efficiency and solution accuracy. Essentially, the study aims to provide insights for CS/SE educators and tool designers on when multimodal GenAI support is beneficial, when it adds overhead, and how to scaffold its use to promote robust learning and practical skill development.