A structured model for identifying the source of synthetic speech performed almost perfectly on the ASVspoof2019-attr-17 benchmark, but the result is limited to known generators and the study’s tested conditions.
An arXiv preprint proposes TCPα, a post-hoc confidence method for music-information-retrieval models. In tests spanning rāga identification, domain shift and ornamentation detection, it reported stronger failure-prediction performance than comparison targets and improved results when low-confidence predictions were rejected.