In plain words: It adds a small sequence-processing module to tiny shared adapter layers, so a big pretrained audio model can adapt to new tasks without retraining its weights. It matched or beat adapter-only baselines on four audio tasks and five speech recognition languages with fewer parameters.
Abstract · MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio
Fine-tuning Transformer-based foundation models has become the dominant strategy for domain adaptation in audio and speech processing. To reduce the computational and memory costs of this process, parameter-efficient transfer learning (PETL) methods have been widely explored. Meanwhile, Mamba, a recent state-space model, has emerged as a promising alternative to Transformers for sequence modeling. In this work, we present MambAdapter, a parameter-efficient transfer learning approach that integrates Mamba into low-rank bottleneck adapters. Our design combines parameter sharing across adapters with the injection of a lightweight Mamba module, enabling more effective modeling of audio features. We demonstrate that MambAdapter matches or outperforms strong PETL baselines on four audio classification tasks and five speech recognition languages, even when operating under reduced parameter budgets.
Salman Hussain Ali, Umberto Cappellazzo, Mirco Ravanelli
arXiv:2606.15638 · eess.AS, cs.SD · submitted Jun 14, 2026
abstract · pdf · html · Accepted to Interspeech 2026. Code available at: https://github.com/salman-ha/MambAdapter