Article · Original
The article text is unavailable in this language; an existing version is shown.
Congrats @smallest_AI on taking #1 for diarization + ASR on @voicearena_ai.
Its a very effective real-world kind of benchmark: > models hear far-field room audio (i.e. microphone is some distance away from the people speaking), but they’re scored against labels built from a mic on every speaker.
> Realistic input, accurate ground truth. That combination is rare.
Many diarization benchmarks either use clean audio or have messy labels. @voicearena_ai avoids both: models get far-field recordings of real in-person conversations, and the ground truth comes from individual mics on each speaker. #1 really means something.
Congrats @smallest_AI.