Abstract
Generating music has a few notable differences from generating images and videos. First, music is an art of the time, necessitating a temporal model. Second, music is usually composed of multiple instruments/tracks with their own temporal dynamics, but collectively they unfold over time interdependently. Lastly, musical notes are often grouped into chords, arpeggios or melodies in polyphonic music, and thereby introducing a chronological ordering of notes is not naturally suitable. In this thesis, we investigate several topics about symbolic multi-track polyphonic music generation under the framework of Conlutional generative adversarial networks (GANs), including controllability, accompaniment, network architecture, temporal modeling. We trained and compared the models on two common formats: lead sheet and band score. To evaluate the generative results, a few intra-track and inter-track objective metrics are also proposed. The integrated survey from data representation, pre-processing, to qualitative evaluation between various architectures offers us more insights about composing music and also re-examining the efficiency and limitation of the deep learning models.