Hybrid Learning Module-Based Transformer for Multitrack Music Generation With Music Theory
Bibliographic Data
| ID | 22106982 |
|---|---|
| Authors | Yun Tie (0000-0002-8258-6206, Zhengzhou University), Xin Guo (0000-0002-7465-9356, Zhengzhou University), Donghui Zhang (0000-0003-3830-4360, Zhengzhou University), Jiessie Tie (0000-0002-3934-4236, University of Toronto), Lin Qi (0000-0003-4005-1702, Zhengzhou University), Yuhang Lu (0009-0007-9394-8587, Zhengzhou University) |
| Year | 2025 |
| Volume | 12 |
| Issue | 2 |
| Pages | 862-872 |
| Publication date | 2025-04-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | IEEE Transactions on Computational Social Systems (JOURNAL) |
| Journal identifiers | ISSN: 2329-924X • E-ISSN: 2373-7476 |
| Publisher | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/tcss.2024.3486604 |
| OpenAlex | W4404520762 |
| Language | EN |
| References cited | 27 |
In recent years, multitrack music generation has garnered significant attention in both academic and industrial spheres for its versatile utilization of various instruments in collaborative settings. The primary challenge lies in achieving a harmonious balance within individual tracks and fostering effective collaboration across multiple tracks. To address this issue, this article introduces a pioneering hybrid learning encoder architecture. Each music track's encoder is implemented as an independent transformer architecture, preserving self-attention mechanisms within a single track and interattention mechanisms between different tracks. The resulting features are then seamlessly integrated into the decoder through concatenation. Of particular significance, previous multitrack music generation efforts have predominantly operated under unconditional settings, yielding music that lacks practical value due to noncompliance with established music theory principles. Recognizing this limitation, the article proposes a novel approach to multitrack music generation guided by music theory rules. Employing reinforcement learning techniques, the decoder-generated music serves as the initial state. Positive feedback is provided when the generated music adheres to music theory rules; conversely, negative feedback is applied to compel the multitrack music to align with widely accepted music theory principles. Finally, comprehensive simulation validation is conducted on both the publicly available LMD dataset and the self-constructed MUT dataset. The plethora of experimental results overwhelmingly corroborates the efficacy of the proposed methodology
Art · Electrical engineering · Electronic engineering · Music theory · Musical · Speech recognition · Transformer · Visual arts · Voltage · Computer Science · Engineering · Music and Audio Processing · Music Technology and Sound Studies · Neuroscience and Music Perception
| Citation velocity | historical |
|---|---|
| Highly cited | No |