M. Heikkinen, A. Politis, K. Drossos, and T. Virtanen, "Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions," in Proceedings of the 19th International Workshop on Acoustic Signal Enhancement (IWAENC), Milan, Italy, 2026
Direction-of-arrival (DoA) estimation is a key component of multichannel audio processing, yet many deep learning approaches remain tied to the microphone arrays used during training and generalize poorly to unseen devices. This paper proposes an array-generic neural DoA estimation framework using measured or simulated complex directional array transfer functions (ATFs) matched to real-world multi-microphone devices. The method processes multichannel spectrograms and ATF metadata with separate convolutional encoders, fuses the resulting representations through cross-attention, and predicts source directions using a multi-source Cartesian vector output formulation. Experiments on simulated 2D and 3D localization tasks under reverberation and diffuse babble noise show that the proposed approach generalizes to previously unseen arrays, including mobile-phone-like configurations, without major performance degradation, while remaining competitive with conventional and learning-based baselines.
This paper addresses a major limitation of neural sound-source localization: most deep-learning methods work well only with the microphone array they were trained on, making deployment to new devices difficult without retraining. The proposed method instead uses detailed acoustic information about the microphone array—its array transfer functions, which describe how sounds arriving from different directions are captured by the microphones—rather than relying only on microphone positions. A neural network separately processes the recorded audio and these array characteristics, then combines them through cross-attention to estimate the directions of one or more sound sources. This allows the same model to operate on previously unseen microphone layouts, including realistic mobile-phone-like arrays where the device body alters the incoming sound field. Experiments in reverberant and noisy environments show that the method generalizes to new arrays with only moderate loss in performance and remains effective as the number of simultaneous sources increases. Classical MUSIC localization is still more accurate for simple single-source cases, but its performance degrades much faster as scenes become more complex, whereas the proposed neural approach remains more stable and can jointly detect and localize multiple sources. The results demonstrate that conditioning neural localization systems on the actual acoustic behaviour of a device is a promising route toward microphone-array-independent spatial audio processing.