We introduce control models for LLM-powered code completion in JetBrains IDEs: ML classifiers which trigger inference and filter the generated suggestions to better align them with users and reduce unnecessary requests. To this end, we evaluate boosting- and transformer-based architectures on an offline dataset of real code completions with $n=98$ users. We further report the classification performance of our boosting-based approach on a range of syntactically diverse languages; and detail how they generalise to a production environment, where they increase inference efficiency by 16% while simultaneously improving completion quality metrics. With this study, we hope to demonstrate the potential in using auxiliary models for smarter in-IDE integration of LLM-driven features, highlight fruitful future directions, and open problems.