正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
@shobhitbanga Whisper Medium sitting at 85% word error rate on Bengali shows how badly the region was underserved,
语音 人工智能 发布 API
自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。
Voice Arena 称,Monsoon ASR 包含逾 10 万小时语音数据,覆盖 50 种语言,多数为使用资源较少的语言。该机构报告称,在 Bengali FLEURS 测试中,用该数据集微调 Whisper Medium 后,词错误率从 85.27% 降至 7.65%;有意授权数据集的机构还可通过模型 API 在内部基准上测试。
该语言的正文暂不可用,当前显示已有版本。
@shobhitbanga Whisper Medium sitting at 85% word error rate on Bengali shows how badly the region was underserved,
Crazy week for Voice Arena at Interspeech in Sydney.
80+ organisations have asked to license Monsoon ASR corpus since we launched it seven days ago.
The most common reason why labs are interested: Monsoon promises results.
When we decided to build datasets at Voice Arena, we set one rule. Either the dataset promises results, or we don't build it.
Monsoon promises results. On Bengali FLEURS, fine-tuning Whisper Medium on Monsoon took its LLM word error rate from 85.27% down to 7.65%. The other thing labs like is that we give them access to our model API. They can test it on their own internal benchmarks and see for themselves whether it will improve their models.
You can see the detailed results and get API access here: https://t.co/m6RSiF9tWP
Monsoon is 100,000 hours across 50 languages from around the world, and most of them are long-tail languages. 100 languages by February, 1,000 by the end of 2027.
It's only the first dataset Voice Arena has launched. We are excited to keep working on the key problems that get us closer to the dream of machines that talk like humans.
在 X 查看回复的帖子
Introducing Monsoon ASR ⚡️ @voicearena_ai
Speech recognition does not have a model problem anymore. It has a data problem.
The best ASR systems are approaching human-level performance in English. But move into the long tail of the world's languages, especially real, conversational speech and error rates can still be 5-10X higher.
Today, we're releasing Monsoon ASR: a new generation of training data built specifically to close that gap.
50 languages. 100,000+ hours. Dense spontaneous speech.
And one goal: Single-digit WER across the world's languages.🧵
在 X 查看上下文