M. Heikkinen, A. Politis, K. Drossos, and T. Virtanen, "Beyond Omnidirectional: Neural Ambisonics Encoding for Arbitrary Microphone Directivity Patterns using Cross-Attention," in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2026
We present a deep neural network approach for encoding microphone array signals into Ambisonics that generalizes to arbitrary microphone array configurations with fixed microphone count but varying locations and frequency-dependent directional characteristics. Unlike previous methods that rely only on array geometry as metadata, our approach uses directional array transfer functions, enabling accurate characterization of real-world arrays. The proposed architecture employs separate encoders for audio and directional responses, combining them through cross-attention mechanisms to generate array-independent spatial audio representations. We evaluate the method on simulated data in two settings: a mobile phone with complex body scattering, and a free-field condition, both with varying numbers of sound sources in reverberant environments. Evaluations demonstrate that our approach outperforms both conventional digital signal processing-based methods and existing deep neural network solutions. Furthermore, using array transfer functions instead of geometry as metadata input improves accuracy on realistic arrays.
This paper tackles a practical problem in spatial audio: how to convert recordings from the small, irregular microphone arrays found in real devices into Ambisonics, a device-independent format used for immersive and 3D audio. Existing neural methods can work well for a specific microphone layout, but often need to be retrained when the array changes, and methods that only use microphone positions cannot describe more realistic effects such as microphone directionality or sound scattering caused by the device body. The proposed method addresses this by giving the neural network not only the recorded audio but also a detailed description of how the microphone array responds to sound coming from different directions and frequencies. Separate neural encoders process the audio and these directional characteristics, and cross-attention combines the two to create an array-independent spatial representation from which the Ambisonics signal is generated. The method is evaluated on simulated microphone arrays in both free-field conditions and a more realistic mobile-phone scenario with complex body scattering. It achieves the best overall signal reconstruction quality in both settings, outperforming conventional Ambisonics encoding and, in realistic phone-like conditions, existing neural approaches. Importantly, the model can generalize to previously unseen microphone configurations without retraining and does so with fewer than one million parameters, showing that detailed knowledge of a device’s directional acoustic behaviour can enable compact and flexible neural spatial-audio capture systems.
M. O. Heikkinen, K. Drosos, A. Politis, and T. Virtanen, “Method and apparatus for training and using a microphone geometry assisted encoder model to generate spatial audio signals,” U.S. Patent US20260065918A1, filed Aug. 28, 2025; published Mar. 05, 2026
A system for training a microphone geometry assisted encoder model and then utilizing the trained model to generate spatial audio signals that have been captured by a plurality of microphones. In a method for generating spatial audio signals, the method includes receiving geometry data related to a plurality of microphones of an audio capturing device and audio signal data captured by the plurality of microphones. The method also includes generating a spatial audio signal based on an output of a trained microphone geometry assisted encoder model. The trained microphone geometry assisted encoder model includes a geometry encoder configured to encode the geometry data and a signal encoder configured to encode the audio signal data. The trained microphone geometry assisted encoder model further includes a signal decoder having a plurality of layers and configured to generate the output upon which the spatial audio signal is based.
This patent addresses a practical limitation of machine-learning-based spatial audio capture: neural models are typically designed for a particular microphone arrangement, so using the same technology on a different phone or recording device may require generating new training data and retraining the model. The proposed approach makes the microphone configuration itself an input to the neural network. A dedicated geometry encoder processes information describing the microphone array—such as the number and positions of the microphones and, in some embodiments, their directional acoustic responses—while a separate encoder processes the recorded audio. The two sources of information are combined by the model, which predicts a set of filters that transforms the microphone recordings into a spatial audio representation such as Ambisonics. The model can be trained using many simulated sound scenes and different microphone configurations so that it learns not only how to reconstruct spatial audio, but also how that reconstruction should change when the microphone arrangement changes. The main benefit is therefore a spatial-audio capture system that can be deployed across devices with different microphone layouts with little or no retraining, reducing the engineering effort normally required to adapt a neural spatial-audio system to each new device.
M. Heikkinen, A. Politis, K. Drossos and T. Virtanen, "Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays," in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025
Using deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained separately for each MA. This paper proposes a DNN-based method for Ambisonics encoding that can generalize to arbitrary MA geometries unseen during training. The method takes as inputs the MA geometry and MA signals and uses a multi-level encoder consisting of separate paths for geometry and signal data, where geometry features inform the signal encoder at each level. The method is validated in simulated anechoic and reverberant conditions with one and two sources. The results indicate improvement over conventional encoding across the whole frequency range for dry scenes, while for reverberant scenes the improvement is frequency-dependent.