Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation [Conference]
Y. Wang, A. Politis, K. Drossos, and T. Virtanen, "Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation," in Proceedings of the IEEE International Workshop on Machine Learning for Signal Processing (MLSP), Atlanta, U.S.A., 2026
Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its challenge lies in predicting cross-segment permutations. This paper proposes a training-free dynamic clustering approach for cross-segment permutation alignment using speaker embedding reference pools. The method predicts the permutation using the cosine similarity between current segment embeddings and the reference pools. The approach updates reference pools by retaining the most representative speaker embeddings based on their overall cosine similarity with existing references. As a plug-and-play post-processing module compatible with existing separation models, the proposed method demonstrates superior performance compared to existing methods on dense and sparse long speech scenarios, particularly in challenging sparse scenarios with extended utterance gaps, and further shows robustness to speaker count estimation errors in unknown speaker count scenarios.