Logo image
Ultra-low-latency distributed deep neural network over hierarchical mobile networks
Conference paper

Ultra-low-latency distributed deep neural network over hierarchical mobile networks

Jen-I Chang, Jian-Jhih Kuo, Chi-Han Lin, Wen-Tsuen Chen and Jang-Ping Sheu
2019 IEEE Global Communications Conference, GLOBECOM 2019 - Proceedings, 9014122
12/2019

Abstract

Deep neural network Early inference Hierarchical mobile network Model deployment Model partition Computer Networks and Communications Hardware and Architecture Information Systems Signal Processing Information Systems and Management Safety Risk Reliability and Quality Media Technology Health Informatics
Recently, the notions of partitioning the Deep Neural Network (DNN) model over the multi-level computing units and making a fast inference with the early- inference technique have been proposed to shorten the inference time. Such computing units form a hierarchical mobile network to provide locality-aware computation, and the early-inference technique allows the prediction results to early exit the model with a probability. However, an inadequate model partition and misapply early inference may prolong response time. Previous studies focus on the classifier design for early inference, and thus, the optimal model partition with classifier deployment has not been explored. In this paper, we study DEMAND-OPE to consider response time and throughput. We first design the COLT for the simplified DEMAND-OPE without Optional Exit Points (DEMAND) to carefully balance the computing time and data transfer time. Then, an extension termed COLT- OPE is developed to achieve the lower response time. Simulation results show that our algorithms (COLT- OPE) outperform previous methods by 200%.

Metrics

1 Record Views

Details

Logo image