正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
Congrats @smallest_AI on taking #1 for diarization + ASR on @voicearena_ai.
Its a very effective real-world kind of benchmark: > models hear far-field room audio (i.e. microphone is some distance away from the people speaking), but they’re scored against labels built from a mic on every speaker.
> Realistic input, accurate ground truth. That combination is rare.
Many diarization benchmarks either use clean audio or have messy labels. @voicearena_ai avoids both: models get far-field recordings of real in-person conversations, and the ground truth comes from individual mics on each speaker. #1 really means something.
Congrats @smallest_AI.