October 20, 2020

AI outperforms humans in speech recognition

by Monika Landgraf, Karlsruhe Institute of Technology

Following a conversation and transcribing it precisely is one of the biggest challenges in artificial intelligence (AI) research. For the first time now, researchers of Karlsruhe Institute of Technology (KIT) have succeeded in developing a computer system that outperforms humans in recognizing such spontaneously spoken language with minimum latency. This is reported on arXiv.org.

"When people talk to each other, there are stops, stutterings, hesitations, such as 'er' or 'hmmm,' laughs and coughs," says Alex Waibel, Professor for Informatics at KIT. "Often, words are pronounced unclearly." This makes it difficult even for people to make accurate notes of a conversation. "And so far, this has been even more difficult for AI." KIT scientists and staff of KITES, a start-up company from KIT, have now programmed a computer system that executes this task better than humans and quicker than other systems.

Waibel already developed an automatic live translator that directly translates university lectures from German or English into the languages spoken by foreign students. This "Lecture Translator" has been used in the lecture halls of KIT since 2012. "Recognition of spontaneous speech is the most important component of this system," Waibel explains, "as errors and delays in recognition make the translation incomprehensible. On conversational speech, the human error rate amounts to about 5.5%. Our system now reaches 5.0%." Apart from precision, however, the speed of the system to produce output is just as important so students can follow the lecture live. The researchers have now succeeded in reducing this latency to one second. This is the smallest reported latency reached by a speech recognition system of this quality to date, says Waibel.

Error rate and latency are measured using the standardized and internationally recognized, scientific "switchboard-benchmark" test. This benchmark (defined by US NIST) is widely used by international AI researchers in their competition to build a machine that comes close to humans in recognizing spontaneous speech under comparable conditions, or even outperforming them.

According to Waibel, fast, high accuracy speech recognition is an essential step for further downstream processing. It enables dialog, translation, and other AI modules to provide better voice based interaction with machines.

More information: Nguyen et al., Super-Human Performance in Online Low-latency Recognition of Conversational Speech. arXiv:2010.03449 [cs.CV]. arxiv.org/abs/2010.03449

Provided by Karlsruhe Institute of Technology

Citation: AI outperforms humans in speech recognition (2020, October 20) retrieved 17 July 2024 from https://techxplore.com/news/2020-10-ai-outperforms-humans-speech-recognition.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.

Explore further

Machine voice recognition reaches human parity

693 shares

Feedback to editors

The magnet trick: New invention makes vibrations disappear

1 hour ago

Creating and verifying stable AI-controlled robotic systems in a rigorous and flexible way

2 hours ago

Unlocking the potential of rust: High-efficiency green hydrogen production from hematite

2 hours ago

Scientists bridge the 'valley of death' in carbon capture technologies

2 hours ago

Flexible electronics researchers develop a completely stretchy lithium-ion battery

6 hours ago

A strategy to enhance the stability of perovskite solar cells under reverse bias conditions

7 hours ago

Engineers evaluate cybersecurity risks associated with EV fast-charging equipment

22 hours ago

Machine learning framework maps global rooftop growth for sustainable energy and urban planning

Jul 16, 2024

Giving drones wrap-and-grip wings to allow them to land on poles and tree limbs

Jul 16, 2024

Large language models make human-like reasoning mistakes, researchers find

Jul 16, 2024

Load comments (0)

AI outperforms humans in speech recognition

The magnet trick: New invention makes vibrations disappear

Creating and verifying stable AI-controlled robotic systems in a rigorous and flexible way

Unlocking the potential of rust: High-efficiency green hydrogen production from hematite

Scientists bridge the 'valley of death' in carbon capture technologies

Flexible electronics researchers develop a completely stretchy lithium-ion battery

A strategy to enhance the stability of perovskite solar cells under reverse bias conditions

Engineers evaluate cybersecurity risks associated with EV fast-charging equipment

Machine learning framework maps global rooftop growth for sustainable energy and urban planning

Giving drones wrap-and-grip wings to allow them to land on poles and tree limbs

Large language models make human-like reasoning mistakes, researchers find

Machine voice recognition reaches human parity

One class in all languages

Google Brain posse takes neural network approach to translation

Facebook unveils machine learning translator for 100 languages

Microsoft claims its new speech recognition system on par with human capabilities

Google introduces real-time extended voice translation

Creating and verifying stable AI-controlled robotic systems in a rigorous and flexible way

New system enables intuitive teleoperation of a robotic manipulator in real-time

Machine learning framework maps global rooftop growth for sustainable energy and urban planning

Microsoft unveils software that allows LLMs to work with spreadsheets

New technique to assess a general-purpose AI model's reliability before it's deployed

Large language models make human-like reasoning mistakes, researchers find

Phys.org

Medical Xpress

Science X

AI outperforms humans in speech recognition

The magnet trick: New invention makes vibrations disappear

Creating and verifying stable AI-controlled robotic systems in a rigorous and flexible way

Unlocking the potential of rust: High-efficiency green hydrogen production from hematite

Scientists bridge the 'valley of death' in carbon capture technologies

Flexible electronics researchers develop a completely stretchy lithium-ion battery

A strategy to enhance the stability of perovskite solar cells under reverse bias conditions

Engineers evaluate cybersecurity risks associated with EV fast-charging equipment

Machine learning framework maps global rooftop growth for sustainable energy and urban planning

Giving drones wrap-and-grip wings to allow them to land on poles and tree limbs

Large language models make human-like reasoning mistakes, researchers find

Related Stories

Machine voice recognition reaches human parity

One class in all languages

Google Brain posse takes neural network approach to translation

Facebook unveils machine learning translator for 100 languages

Microsoft claims its new speech recognition system on par with human capabilities

Google introduces real-time extended voice translation

Recommended for you

Creating and verifying stable AI-controlled robotic systems in a rigorous and flexible way

New system enables intuitive teleoperation of a robotic manipulator in real-time

Machine learning framework maps global rooftop growth for sustainable energy and urban planning

Microsoft unveils software that allows LLMs to work with spreadsheets

New technique to assess a general-purpose AI model's reliability before it's deployed

Large language models make human-like reasoning mistakes, researchers find

Your Privacy