Google’s DeepMind develops creepy, ultra-realistic human speech synthesis | Science! | Geek.com

A speech waveform, with a zoomed section below showing the individual audio samples as dots.

Most text-to-speech systems in 2016 worked by stitching together short fragments recorded from one speaker, or by passing a model’s output through signal processing algorithms called vocoders. DeepMind’s WaveNet instead models the raw audio waveform directly, one sample at a time.

Raw audio is hard to model because it typically contains 16,000 or more samples per second, with structure at many time scales. WaveNet is a convolutional neural network whose layers use different dilation factors, so it can take thousands of previous samples into account when predicting the next one.

In blind listening tests on US English and Mandarin Chinese, WaveNet cut the gap between Google’s best existing systems and human speech by over 50%. Because it models raw audio, the same network can also generate other sounds, including piano music.

via WaveNet: A generative model for raw audio | Google DeepMind

Monday September 12th 2016