PhD Seminar: Automatic Recognition of Disordered Speech: Domain Generalization Approaches for Inclusive ASR

Monday, September 14, 2026 10:00 am - 11:00 am EDT (GMT -04:00)

Candidate: Nada Gohider
Date: September 14, 2026
Time: 10:00 AM
Location: Online
Supervisor: Otman Basir

All are welcome!

Abstract:

Automatic Speech Recognition (ASR) systems have achieved remarkable performance for typical speech; however, their accuracy degrades substantially when applied to speech produced by individuals with speech disorders. This performance gap is primarily attributed to the limited availability of labelled disordered speech and the significant domain shift arising from the large variability in the characteristics of impaired speech across speakers, disorders, and severity levels. Consequently, developing robust ASR systems that generalize effectively to previously unseen disordered speech domains remains a fundamental challenge.

The current research addresses this challenge by proposing a domain generalization framework for robust disordered speech recognition based on three complementary learning paradigms: self-supervised learning (SSL), multi-task learning (MTL), and meta-learning. Rather than treating these paradigms as independent solutions, the proposed framework considers them as complementary building blocks toward a conceptually unified framework where each component addresses a different aspect of the domain generalization problem. SSL is investigated as a representation learning framework that exploits large-scale unlabelled speech to learn robust and transferable acoustic representations, thereby reducing reliance on the limited labelled disordered speech available for training. Building upon these representations, MTL is employed to improve the primary ASR task through auxiliary supervision that explicitly regulates the contribution of domain-specific information while preserving linguistically relevant features. Both cooperative and adversarial MTL frameworks are investigated to evaluate different strategies for learning shared representations across the training tasks. Finally, meta-learning is explored as a mechanism for learning transferable meta-knowledge that enables rapid adaptation to previously unseen speech domains, thereby enhancing the model’s ability to generalize across diverse speech disorders.

Comprehensive experiments are conducted to evaluate each of the proposed frameworks under a range of settings and examined factors. The experimental evaluation demonstrates that, although supervised speech foundation models generally improve ASR performance, large-scale SSL pretrained models exhibit substantially stronger generalization to previously unseen domains, highlighting their ability to mitigate the domain shift problem inherent in disordered speech recognition. This observation motivates the adoption of SSL as the representation learning foundation of the proposed framework. Furthermore, among the investigated MTL approaches, adversarial MTL consistently provides the greatest robustness by learning domain-invariant representations that improve generalization without suffering from the negative transfer observed in conventional cooperative MTL. In contrast, while cooperative MTL does not consistently outperform single-task learning when sufficient labelled speech is available, it demonstrates clear advantages under low-resource conditions, where auxiliary supervision promotes the learning of more informative shared representations. Finally, the proposed parameter-efficient fine-tuning (PEFT)-aware Model-Agnostic Meta-Learning (MAML) based meta-learning framework further enhances robustness under limited-resource scenarios by improving adaptation to previously unseen speech domains. Overall, this thesis demonstrates that self-supervised representation learning, adversarial auxiliary supervision, and meta-learning provide complementary mechanisms for addressing the domain shift problem in disordered speech recognition. The proposed framework establishes a principled approach for learning robust speech representations, contributing to the development of more reliable and accessible ASR technologies for individuals with speech impairments.