In plain words: This model lets a transformer handle complex numbers—pairs that capture a wave's size and timing, as sound and radio signals are—by rewriting its core steps to use both parts. It beat the previous best results on music and radio-signal tasks.
Abstract
While deep learning has received a surge of interest in a variety of fields in recent years, major deep learning models barely use complex numbers. However, speech, signal and audio data are naturally complex-valued after Fourier Transform, and studies have shown a potentially richer representation of complex nets. In this paper, we propose a Complex Transformer, which incorporates the transformer model as a backbone for sequence modeling; we also develop attention and encoder-decoder network operating for complex input. The model achieves state-of-the-art performance on the MusicNet dataset and an In-phase Quadrature (IQ) signal dataset.
Muqiao Yang, Martin Q. Ma, Dongyu Li, Yao-Hung Hubert Tsai, Ruslan Salakhutdinov
arXiv:1910.10202 · cs.LG, cs.SD, eess.AS, stat.ML · submitted Oct 22, 2019 · updated Aug 6, 2021
abstract · pdf · html