In plain words: They estimate tongue and jaw movement from the sound, then score voice-converted speech by how far those movements drift from natural speech. Converted speech showed more movement errors and less shared information, and this score matched human ratings better than the standard spectral-distance score.
Abstract
We propose a novel application based on acoustic-to-articulatory inversion towards quality assessment of voice converted speech. The ability of humans to speak effortlessly requires coordinated movements of various articulators, muscles, etc. This effortless movement contributes towards naturalness, intelligibility and speakers identity which is partially present in voice converted speech. Hence, during voice conversion, the information related to speech production is lost. In this paper, this loss is quantified for male voice, by showing increase in RMSE error for voice converted speech followed by showing decrease in mutual information. Similar results are obtained in case of female voice. This observation is extended by showing that articulatory features can be used as an objective measure. The effectiveness of proposed measure over MCD is illustrated by comparing their correlation with Mean Opinion Score.
Avni Rajpal, Nirmesh J. Shah, Mohammadi Zaki, Hemant A. Patil
arXiv:1511.04867 · cs.SD · submitted Nov 16, 2015 · updated Nov 23, 2015
abstract · pdf · The paper is withdrawn from the arxiv. Author doesnot want circulation of unpublished unverified results