资讯
Hugging Face Blog
AI
LLM
Dataset
Platform
中文标题
推出 Real World VoiceEQ:评估语音 AI 的人类感知质量
English Title
Introducing Real World VoiceEQ: Measuring the human quality of voice AI
David Ayllon, Alice, Jeff Brooks, Franc Camps Febrer, Jakub Piotr Cłapa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran, Rashish Tandon, Panagiotis Tzirakis
发布时间
2026/7/15 08:00:00
来源类型
blog
语言
en
摘要
中文对照

语音正迅速成为 AI 的主要交互界面。从客户服务、医疗健康到教育、娱乐及个人助理,语音正日益取代文本,成为人们与 AI 交互的主要方式。过去几年间,语音模型性能显著提升:词错误率持续下降,延迟已达到可支持自然对话的水平,许多既有基准测试也正趋于饱和。然而,任何经常使用语音 AI 的人都能察觉,某些方面仍显异常。

English Original

Voice is rapidly becoming AI's primary interface. From customer support and healthcare to education, entertainment, and personal assistants, speech is increasingly replacing text as the way people interact with AI. Over the last few years, voice models have improved dramatically. Word error rates continue to fall, latency has reached conversational speeds, and many established benchmarks are approaching saturation. Yet anyone who regularly uses voice AI knows something still feels off.

正文
中文全文

语音正迅速成为人工智能的主要交互界面。从客户服务与医疗健康,到教育、娱乐及个人助理,语音正日益取代文本,成为人们与人工智能互动的方式。过去几年间,语音模型取得了显著进步:词错误率持续下降,延迟已达到可支持自然对话的水平,许多既定基准测试指标也正趋近饱和。然而,任何经常使用语音人工智能的人都能察觉,某些方面仍不尽如人意。为量化这些特质,我们构建了Real World VoiceEQ——一项专为评估语音交互人类化水平而设计的基准测试。它衡量语音系统能否识别、生成并响应文本转录所遗漏的声学信息,包括语调、情绪、说话人身份以及背景环境等。 Real World VoiceEQ基于逾100万条来自不同人口统计特征、说话风格及声学环境的个体人类评分构建而成。当前基准包含78.5万条文本转语音(TTS)评分和4.8万条语音转文本(STS)评分,是迄今规模最大的语音人工智能人类评估之一。所有评估均通过Kairos——我们自主研发的灵活、原生语音评估平台——完成。同一套基础设施亦赋能前沿人工智能实验室与企业开展定制化评估,适配特定应用场景;识别生产环境中语音系统的细粒度失效模式;生成人类偏好数据;并通过强化学习与人类反馈持续优化模型。 追求单一“最优”语音模型的竞争,正让位于对一系列专业化能力的构建。随着语音人工智能日趋成熟,衡量其进展愈发需要独立评估各项能力,而非将其统合为一个总体得分。在我们的TTS评估中,没有任何系统配置能在全部八个能力维度中均跻身前五——这进一步印证了并不存在一个普适意义上的“最佳”语音模型。设想一位银行客服代理询问您是否认出一笔潜在欺诈交易:“是的”这一自信答复与“……是的……”这一迟疑回应,尽管文本转录完全相同,却可能传达截然不同的含义。人类可即刻识别这种差异,而当今许多语音模型尚无法做到。 大语言模型(LLM)目前已广泛用于文本模型评估,但我们的研究结果表明,语音语言模型(SLM)在语音评估中需更审慎地使用。当我们对比领先SLM与经训练的人类评分员在文本转语音评估任务中的表现时发现:在发音准确性等具有明确、可验证答案的任务上,二者一致性最高;而在更具主观性的评估中,一致性则明显下降。SLM有时会仅依据文本上下文线索推断情绪,而在诸如“语音是否契合某一表演角色”或“语音身份是否保持一致”等开放式判断任务中,其一致性最弱。自动化评估器对于定义清晰的任务颇具价值,但在依赖声学语境、感知判断与社会性解读的评估中,尚无法替代人类听者。 随着语音成为人工智能最具标志性的交互界面之一,仅凭速度与技术准确率已不足以决定哪些系统最终胜出。用户最终选择的,将是那些能在真实世界复杂对话中——而不仅限于理想化基准测试条件下——真正理解、表达并回应人类意图的模型。 数十年来,语音人工智能的发展始终围绕标准化基准上的量化指标展开:从衡量转录准确率的词错误率(WER),到评估语音质量的客观感知指标(如PESQ与DNSMOS)。我们希望Real World VoiceEQ能够拓展这一范式,提供一种以人类体验为根基的评估指标,用以衡量合成语音交互各构成要素的表现。 阅读完整技术报告,浏览公开排行榜;或联系我们,了解Hume如何基于Real World VoiceEQ评估您的语音模型或智能体,或为您特定应用场景定制评估方案。

English Original

Voice is rapidly becoming AI's primary interface. From customer support and healthcare to education, entertainment, and personal assistants, speech is increasingly replacing text as the way people interact with AI. Over the last few years, voice models have improved dramatically. Word error rates continue to fall, latency has reached conversational speeds, and many established benchmarks are approaching saturation. Yet anyone who regularly uses voice AI knows something still feels off. To measure those qualities, we built Real World VoiceEQ—a benchmark designed to evaluate the human quality of voice interaction. It assesses whether voice systems can recognize, produce, and respond to the acoustic information transcripts leave out, from tone and emotion to speaker identity and background context. Real World VoiceEQ was developed from more than 1 million individual human ratings collected across different demographics, speaking styles, and acoustic environments. The current benchmark includes 785,000 TTS ratings and 48,000 STS ratings, making it one of the largest human evaluations of voice AI conducted to date. Every evaluation was conducted using Kairos, our flexible, voice-native evaluation platform. The same infrastructure enables frontier AI labs and enterprises to run custom evaluations tailored to specific use cases, identify granular failure modes in production voice systems, generate human preference data, and continuously improve models through reinforcement learning and human feedback. The race for a single "best" voice model is giving way to a collection of specialized capabilities. As voice AI matures, measuring progress increasingly requires evaluating these capabilities independently rather than collapsing them into a single overall score. In our TTS evaluations, no system configuration ranked among the top five across all eight capability groups—underscoring why there is no single "best" voice model. Imagine a banking agent asking whether you recognize a potentially fraudulent transaction. A confident "Yes" and a hesitant "…yes…" may have completely different meanings, even though the transcript is identical. Humans recognize that difference immediately. Many of today's voice models do not. LLMs are now widely used to evaluate text-based models, but our findings suggest that speech-language models (SLMs) should be used more carefully for voice evaluation. When we compared leading SLMs with trained human raters on text-to-speech assessments, agreement was highest on tasks with clear, verifiable answers, such as pronunciation accuracy. Agreement declined on more subjective evaluations. SLMs sometimes appeared to infer emotion from text-based contextual cues, and agreement was weakest for open-ended judgments such as whether a voice fit an acting role or maintained a consistent identity. Automated evaluators can be valuable for well-defined tasks, but they are not yet a substitute for human listeners when judgments depend on acoustic-context, perception, and social interpretation. As voice becomes one of AI's defining interfaces, speed and technical accuracy alone will no longer determine which systems succeed. The models people ultimately choose will be those that can understand, express, and respond like humans—not just under ideal benchmark conditions, but across the complexity of real-world conversation. For decades, speech AI has advanced by optimizing against quantitative metrics on standardized benchmarks; from WER for transcription accuracy to objective perceptual metrics like PESQ and DNSMOS for speech quality. We hope Real World VoiceEQ can extend this paradigm by providing a human-grounded metric for evaluating the components of synthetic voice interactions. Read the full technical report and explore the public leaderboards—or get in touch to learn how Hume can evaluate your voice model or agent using Real World VoiceEQ, or design custom evaluations tailored to your specific use case.

元数据
来源Hugging Face Blog
类型资讯
抽取状态raw
关键词
AI
LLM
Dataset
Platform