Controller synthesis is a noted approach that enables automatically generating control strategies that satisfy given specifications. To enhance efficiency, Directed Controller Synthesis prunes the search space by incrementally constructing a partial view of the system, aiming to find a valid controller without exhaustive exploration. This process is steered by an exploration policy (i.e., heuristic), and Reinforcement Learning has proven highly effective for learning such policies. However, a key challenge is anisotropic generalization, i.e., a single policy trained on specific domain parameters is specialized, performing well in certain scenarios while remaining fragile in others. To this end, we propose a Mixture-of-Experts framework that synergistically combines multiple policies, leveraging their complementary strengths to form a more robust exploration policy. The evaluation of the Air Traffic benchmark shows that our proposal significantly increases the number of solvable instances, with a modest increase in computational overhead.
Maria Casimiro INESC-ID, IST, University of Lisbon & S3D, Carnegie Mellon University, Valentim Romão INESC-ID, Instituto Superior Técnico, Universidade de Lisboa, Paolo Romano University of Lisbon, Portugal, Luis Rodrigues INESC-ID, IST, ULisboa, David Garlan Carnegie Mellon University