摘要
The notion of "noise" in musical contexts could be misleading because, when appropriately edited and arranged, noise can sound pleasant to human ears. In this study, we developed a system that transforms Japanese pop songs to the chiptune style. A neural network model first unmixed the original music into four tracks: vocals, bass, percussion, and residual. Afterwards, we utilized non-negative matrix factorization deconvolution (NMFD) to analyze the percussion track and detect the beats of the bass drum, the snare drum, the hi-hat, and the cymbal separately. The NMFD template length for the hi-hat and the cymbal was set three times longer than that of the drums to capture their difference in decay time. To render the content into the chiptune style, different channels of filtered noise were selected for the four drumkit instruments, while the vocals and the bass were re-generated using square waves and triangular waves, respectively. Evaluation by 35 listeners showed that the perceived overall naturalness ranged between 6.0 and 7.9 out of 10, and the naturalness of the synthesized percussion was rated 8.9/10. These indicate that listeners can adapt to appreciate music in the lo-fidelity chiptune style which contains automatic replacement of drumkit sounds by filtered noise.