Edge intelligence enables deep learning (DL) inference services to run close to end users, which is essential for latency-sensitive applications such as autonomous driving and smart surveillance. However, satisfying strict latency-oriented service level objectives (SLOs) at the edge remains challenging because inference performance is jointly influenced by the selected DL model variant, the CPU parallelism configuration, and resource interference on shared edge servers. Existing approaches have not adequately considered these three factors in a unified manner. In this paper, we address this gap by proposing a semi-supervised framework that jointly models these three dimensions to predict P90 latency and support serving decisions. Our approach represents candidate DL models through computational graphs structured as directed acyclic graphs (DAGs) and combines unsupervised structural pretraining with supervised runtime-aware graph learning. The framework is integrated into an inference-serving pipeline that filters candidate models according to user accuracy requirements, evaluates feasible thread configurations under current interference conditions, and selects the model and CPU parallelism setting that satisfy the target P90 latency constraint.
Semi-supervised approach for inference serving at the edge
IWCMC 2026, 22nd International Wireless Communications and Mobile Computing Conference, 1-6 June 2026, Wuzhou, China
Type:
Conference
City:
Wuzhou
Date:
2026-06-01
Department:
Communication systems
Eurecom Ref:
8857
Copyright:
© 2026 IEEE. Personal use of this material is permitted. However, permission to reprint/republish this material for advertising or promotional purposes or for creating new collective works for resale or redistribution to servers or lists, or to reuse any copyrighted component of this work in other works must be obtained from the IEEE.
See also:
PERMALINK : https://www.eurecom.fr/publication/8857