Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Hybrid Learning Module-Based Transformer for Multitrack Music Generation With Music Theory

Bibliographic Data

ID22106982
AuthorsYun Tie (0000-0002-8258-6206, Zhengzhou University), Xin Guo (0000-0002-7465-9356, Zhengzhou University), Donghui Zhang (0000-0003-3830-4360, Zhengzhou University), Jiessie Tie (0000-0002-3934-4236, University of Toronto), Lin Qi (0000-0003-4005-1702, Zhengzhou University), Yuhang Lu (0009-0007-9394-8587, Zhengzhou University)
Year2025
Volume12
Issue2
Pages862-872
Publication date2025-04-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueIEEE Transactions on Computational Social Systems (JOURNAL)
Journal identifiersISSN: 2329-924X • E-ISSN: 2373-7476
PublisherInstitute of Electrical and Electronics Engineers (IEEE) (PUBLISHER)
DOI10.1109/tcss.2024.3486604
OpenAlexW4404520762
LanguageEN
References cited27

In recent years, multitrack music generation has garnered significant attention in both academic and industrial spheres for its versatile utilization of various instruments in collaborative settings. The primary challenge lies in achieving a harmonious balance within individual tracks and fostering effective collaboration across multiple tracks. To address this issue, this article introduces a pioneering hybrid learning encoder architecture. Each music track's encoder is implemented as an independent transformer architecture, preserving self-attention mechanisms within a single track and interattention mechanisms between different tracks. The resulting features are then seamlessly integrated into the decoder through concatenation. Of particular significance, previous multitrack music generation efforts have predominantly operated under unconditional settings, yielding music that lacks practical value due to noncompliance with established music theory principles. Recognizing this limitation, the article proposes a novel approach to multitrack music generation guided by music theory rules. Employing reinforcement learning techniques, the decoder-generated music serves as the initial state. Positive feedback is provided when the generated music adheres to music theory rules; conversely, negative feedback is applied to compel the multitrack music to align with widely accepted music theory principles. Finally, comprehensive simulation validation is conducted on both the publicly available LMD dataset and the self-constructed MUT dataset. The plethora of experimental results overwhelmingly corroborates the efficacy of the proposed methodology

Art · Electrical engineering · Electronic engineering · Music theory · Musical · Speech recognition · Transformer · Visual arts · Voltage · Computer Science · Engineering · Music and Audio Processing · Music Technology and Sound Studies · Neuroscience and Music Perception

Citation velocityhistorical
Highly citedNo

Tools

Open DOI
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae