The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Jul. 04, 2017

Filed:

Jul. 08, 2016
Applicant:

Google Inc., Mountain View, CA (US);

Inventors:

Tara N. Sainath, Jersey City, NJ (US);

Ron J. Weiss, New York, NY (US);

Kevin William Wilson, Cambridge, MA (US);

Andrew W. Senior, New York, NY (US);

Arun Narayanan, Santa Clara, CA (US);

Yedid Hoshen, Jerusalem, IL;

Michiel A. U. Bacchiani, Summit, NJ (US);

Assignee:

Google Inc., Mountain View, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G10L 15/16 (2006.01); G10L 15/06 (2013.01); G10L 21/0216 (2013.01); G10L 15/02 (2006.01);
U.S. Cl.
CPC ...
G10L 15/16 (2013.01); G10L 15/02 (2013.01); G10L 15/063 (2013.01); G10L 2021/02166 (2013.01);
Abstract

Methods, including computer programs encoded on a computer storage medium, for enhancing the processing of audio waveforms for speech recognition using various neural network processing techniques. In one aspect, a method includes: receiving multiple channels of audio data corresponding to an utterance; convolving each of multiple filters, in a time domain, with each of the multiple channels of audio waveform data to generate convolution outputs, wherein the multiple filters have parameters that have been learned during a training process that jointly trains the multiple filters and trains a deep neural network as an acoustic model; combining, for each of the multiple filters, the convolution outputs for the filter for the multiple channels of audio waveform data; inputting the combined convolution outputs to the deep neural network trained jointly with the multiple filters; and providing a transcription for the utterance that is determined.


Find Patent Forward Citations

Loading…