Synergizing Zero-Shot Cross-Lingual Alzheimer Detection with Language-Invariant Multimodal Bi-Geometric Adversarial Learning
* Equal contribution.
Research record
TLDR — verified methodology and contribution summary
ORBIT combines cross-attentive multimodal fusion, language adversaries, and spherical-hyperbolic learning for zero-shot cross-lingual Alzheimer detection.
Abstract
In this work, we study zero-shot cross-lingual speech-based Alzheimer’s disease detection (SADD). We hypothesize that learning language-invariant multimodal representations by fusing multilingual speech and text pretrained models is essential for reliable transfer to unseen languages, while adversarial learning suppresses language-specific confounds. Empirical results in zero-shot cross-lingual evaluation show that multimodal fusion consistently outperforms unimodal baselines. We propose ORBIT, a framework combining cross-attentive fusion, multi-tap language adversaries, and complementary spherical-hyperbolic geometric learning with consensus clustering. Across settings, ORBIT achieves the strongest performance compared with unimodal models and simple concatenation-based fusion baselines. This work is recorded as accepted to Interspeech 2026; its arXiv record remains the accessible source.