Microsoft’s Speech Recognition Tech Runs NVIDIA Hardware

Earlier in the week, we reported that Microsoft researchers had set a world record for speech recognition, using a technology it announced this week with GPU-accelerated deep learning to recognise words in a conversation as well as a person does.

The team achieved an error rate of 5.9 percent — the lowest ever for machine transcription — and about as accurate as people who transcribed the same conversation. It’s also a 6 percent improvement over a record Microsoft set only a month ago.

Powerful GPUs = Faster Progress

NVIDIA GPUs and Microsoft’s Cognitive Toolkit (previously known as CNTK), an open source deep learning framework, played key roles in reaching human parity for conversational speech recognition. The cognitive toolkit, which Microsoft announced this week, is a system for deep learning that is used to speed advances in areas such as speech and image recognition and search relevance on GPUs.

By using NVIDIA’s Tesla M40 GPUs, researchers reduced the training time for some language models from months to weeks.

Comment what you think!