"Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens,"in Proc. ICML, 2025. (NVIDIA) [Pub.] [HuggingFace] [Code]



๐Ÿ‘๏ธโ€๐Ÿ—จ๏ธ 1. Introduction / 2. Related works

[โ“โš ๏ธ Research Question]

โ€”SD+ multi-speaker ASR ์˜ integration โ€”**Permutation (**speaker-label alignment)

[๐Ÿ“Ž Related Work] Limitations of previous methods

[๐Ÿ”ฌ Experiment] Real-word dataset (~4spk)

๐Ÿ—ผ 3. Proposed Approach: SortFormer

โš ๏ธ 3.1 Permutation Problem in Diarization

3.2 Diarization Model as a Multi-label Binary Classifier

๐Ÿ–๏ธ3.3 Loss Calculation โ€” Arrival-Time Sorting

3.4 Transformer Encoder Learns to Sort โ€” PE

๐Ÿ—ผ4. Bridging Timestamps and Tokens

๐Ÿงช 5. Experimental Results