Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
* Equal contribution.
Research record
TLDR — verified methodology and contribution summary
MiCuNet uses mixed-curvature fusion and temporal gating to trace emotion, manipulation, and generator source across English and Chinese synthetic speech.
Abstract
In this work, we address the problem of fine-grained traceback of emotional and manipulation characteristics from synthetically manipulated speech. We hypothesize that combining semantic-prosodic cues captured by Speech Foundation Models (SFMs) with fine-grained spectral dynamics from auditory representations can enable more precise tracing of both emotion and manipulation source. To validate this, we introduce MiCuNet, a multitask framework for fine-grained tracing of emotional and manipulation attributes in synthetically generated speech. The approach integrates SFM embeddings with spectrogram-based auditory features through a mixed-curvature projection mechanism spanning Hyperbolic, Euclidean, and Spherical spaces, guided by learnable temporal gating. It simultaneously predicts original emotions, manipulated emotions, and manipulation sources on the EmoFake dataset across English and Chinese subsets. MiCuNet yields consistent improvements over conventional fusion strategies.