Real-time Adapting Routing (RAR): Improving Efficiency Through Continuous Learning in Software Powered by Layered Foundation Models (ICSE 2025 - Software Engineering in Practice (SEIP))

Who

Kirill Vasilevski, Dayi Lin, Ahmed E. Hassan

Track

ICSE 2025 SE In Practice (SEIP)

Time Zone

The program is currently displayed in (GMT-04:00) Eastern Time (US & Canada).

Use conference time zone: (GMT-04:00) Eastern Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Wed 30 Apr 2025 12:15 - 12:30 at 215 - SE for AI 1 Chair(s): Houari Sahraoui

Abstract

To balance the quality and inference cost of a Foundation Model (FM, such as large language models (LLMs)) powered software, people often opt to train a routing model that routes requests to FMs with different sizes and capabilities. Existing routing models rely on learning the optimal routing decision from carefully curated data, require complex computations to be updated, and do not consider the potential evolution of weaker FMs. In this paper, we propose Real-time Adaptive Routing (RAR), an approach to continuously adapt FM routing decisions while using guided in-context learning to enhance the capabilities of weaker FM. The goal is to reduce reliance on stronger, more expensive FMs. We evaluate our approach on different subsets of the popular MMLU benchmark. Our approach routes 50.2% fewer requests to computationally expensive models while maintaining around 90.5% of the general response quality. In addition, the generated guidance from stronger models has shown intra-domain generalization and led to a better quality of responses compared to an equivalent approach with a standalone weaker FM.

Link to Preprint

https://arxiv.org/abs/2411.09837

File attachments

rar_realtime_adapting_routing_icse2025-seip-p121 (rar_realtime_adapting_routing_icse2025-seip-p121.pdf)	2.10MiB

Kirill Vasilevski

Huawei Canada

Dayi Lin

Centre for Software Excellence, Huawei Canada

Canada

Ahmed E. Hassan

Queen’s University