← Back to publications

Synergizing Zero-Shot Cross-Lingual Alzheimer Detection with Language-Invariant Multimodal Bi-Geometric Adversarial Learning

  1. Girish*
  2. Mohd Mujtaba Akhtar*
  3. Farhan Sheth*
  4. Muskaan Singh
  5. Juliana Gerard
  6. Paula McClean
  7. Kongfatt Wong-Lin

* Equal contribution.

Research record

TLDR — verified methodology and contribution summary

ORBIT combines cross-attentive multimodal fusion, language adversaries, and spherical-hyperbolic learning for zero-shot cross-lingual Alzheimer detection.

Abstract

In this work, we study zero-shot cross-lingual speech-based Alzheimer’s disease detection (SADD). We hypothesize that learning language-invariant multimodal representations by fusing multilingual speech and text pretrained models is essential for reliable transfer to unseen languages, while adversarial learning suppresses language-specific confounds. Empirical results in zero-shot cross-lingual evaluation show that multimodal fusion consistently outperforms unimodal baselines. We propose ORBIT, a framework combining cross-attentive fusion, multi-tap language adversaries, and complementary spherical-hyperbolic geometric learning with consensus clustering. Across settings, ORBIT achieves the strongest performance compared with unimodal models and simple concatenation-based fusion baselines. This work is recorded as accepted to Interspeech 2026; its arXiv record remains the accessible source.