欢迎访问中国科学院大学学报,今天是

中国科学院大学学报 ›› 2026, Vol. 43 ›› Issue (4): 566-575.DOI: 10.7523/j.ucas.2024.040

• 电子信息与计算机科学 • 上一篇    

面向丰富曲调要素的影视配乐生成模型

赵冰爽, 罗铁坚(), 王承杰   

  1. 中国科学院大学计算机科学与技术学院,北京 101408
  • 收稿日期:2024-03-04 修回日期:2024-05-06 发布日期:2024-05-29
  • 通讯作者: 罗铁坚
  • 基金资助:
    中国科学院战略先导项目(E0421104)

BM-Transformer: a generative model for film and television soundtracks enriched with melodic elements

Bingshuang ZHAO, Tiejian LUO(), Chengjie WANG   

  1. School of Computer Science and Technology,University of Chinese Academy of Sciences,Beijing 101408,China
  • Received:2024-03-04 Revised:2024-05-06 Published:2024-05-29
  • Contact: Tiejian LUO

摘要:

多模态模型在生成语言、视频和乐曲等任务上表现出极大的潜力,然而在面向丰富曲调要素的背景音乐生成任务上仍面临着情感一致性和专业引导等问题。本文提出多维交互引导和时间比例编码,在音乐表达复合词中嵌入情感标签、韵律密度和韵律强度,生成具有多维交互特性的音乐向量表示。设计具备视频与音乐的节奏和谐对应的影视配乐生成模型,并给出相应的非配对数据驱动的模型网络训练方法。提出的“创作者-观众”双视角评估策略,可以更全面地评估模型的交互性和配乐效果。实验结果表明,该模型在客观评价指标上提高20.6%,推断效率由0.125提升至14.370,并且在主观评价指标上也有明显优势。

关键词: 深度神经网络, 情感一致性, 影视配乐生成模型, 多维交互引导

Abstract:

Multimodal models show great potential in tasks such as generating language, video, and musical scores. However, they still face problems such as emotional consistency and professional guidance for background music generation tasks oriented toward rich melodic elements. In this paper, we propose multidimensional interaction guidance and temporal scaling encoding to generate music vector representations with multidimensional interaction properties by embedding emotion labels, rhythmic density, and rhythmic intensity into music expression compound words. We design the background music transformer (BMT), a model for generating film and television soundtracks with rhythmic and harmonic correspondence between video and music, and provide a corresponding unpaired data-driven network training method for the model. We propose a “creator-audience” dual-view evaluation strategy to evaluate the interactivity and soundtrack effect of the model in a more comprehensive way. The experimental results show that the BMT model improves the objective evaluation index by 20.6%, and the model inference efficiency increases from 0.125 to 14.370, and also has obvious advantages in subjective evaluation index.

Key words: deep neural network, emotional coherence, video soundtrack generation model, multidimensional interaction guidance

中图分类号: