Abstract
Device-to-device (D2D) communication is one of the promising solutions to improve spectrum efficiency and alleviate the mobile traffic explosion. However, interference mitigation and resource allocation in the underlying cellular network is a challenging task. In this paper, we propose a distributed deep reinforcement learning (DRL) based scheme to solve the interference mitigation and resource allocation problem. According to the channel status, each cellular user (CU) and D2D transmitter (D2D TX) will determine the appropriate reused channel and transmit power to maximize the system throughput. We propose a distributed DRL scheme and integrate two hotbooting algorithms into the scheme to improve the system throughput at the early stage of training. Simulation results show that the proposed distributed DRL with hotbooting outperforms the baselines regarding running time, message overhead, and throughput.