摘要
Generative Large Language Models have shown impressive in-context learning abilities, performing well across various tasks with just a prompt. Previous melody-to-lyric research has been limited by scarce high-quality aligned data and un-clear standard for creativeness. Most efforts focused on gen-eral themes or emotions, which are less valuable given cur-rent language model capabilities. In tonal contour languages like Mandarin, pitch contours are influenced by both melody and tone, leading to variations in lyric-melody fit. Our study, validated by the Mpop600 dataset, confirms that lyricists and melody writers consider this fit during their composition pro-cess. In this research, we developed a multi-agent system that decomposes the melody-to-lyric task into sub-tasks, with each agent controlling rhyme, syllable count, lyric-melody align-ment, and consistency. Listening tests were conducted via a diffusion-based singing voice synthesizer to evaluate the qual-ity of lyrics generated by different agent groups.