摘要
The scarcity of expansive datasets for singing quality assessment makes the utilization of complex deep learning methods a considerable challenge. This research presents a method to improve the singing quality prediction based on the feedback from subjective human perception opinion that is learned by the transfer learning methods of self-supervised learning (SSL) speech models. In combination with the CRNN_PH model as the baseline model, the SSL models are integrated into two distinct major architectures: one directly draws features from the pre-trained SSL model (CRNN_PH+SSL), and the other employs the weighted sum (WS) of the output features from different transformer blocks in the SSL model (CRNN_PH+SSL_WS). We conducted comparative experiments on pre-trained SSL models, five on wav2vec 2.0 (W2V2) and two on HuBERT, which were trained over various datasets. It turns out that CRNN_PH+W2V2_base_WS is improved the most on singing quality score prediction that is closely aligning with subjective human perceptions in terms of correlation coefficients and MSE with respect to the ground truth.
| 原文 | 英語 |
|---|---|
| 主出版物標題 | Proceedings of the 5th ACM International Conference on Multimedia in Asia, MMAsia 2023 |
| 發行者 | Association for Computing Machinery, Inc |
| ISBN(電子) | 9798400702051 |
| DOIs | |
| 出版狀態 | 已出版 - 06 12 2023 |
| 事件 | 5th ACM International Conference on Multimedia in Asia, MMAsia 2023 - Hybrid, Tainan, 台灣 持續時間: 06 12 2023 → 08 12 2023 |
出版系列
| 名字 | Proceedings of the 5th ACM International Conference on Multimedia in Asia, MMAsia 2023 |
|---|
Conference
| Conference | 5th ACM International Conference on Multimedia in Asia, MMAsia 2023 |
|---|---|
| 國家/地區 | 台灣 |
| 城市 | Hybrid, Tainan |
| 期間 | 06/12/23 → 08/12/23 |
文獻附註
Publisher Copyright:© 2023 Copyright held by the owner/author(s).
指紋
深入研究「Improve Singing Quality Prediction Using Self-supervised Transfer Learning and Human Perception Feedback」主題。共同形成了獨特的指紋。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver