The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Jan. 06, 2026

Filed:

Feb. 13, 2023
Applicant:

Google Llc, Mountain View, CA (US);

Inventors:

Andrew M. Rosenberg, Brooklyn, NY (US);

Zhehuai Chen, Edgewater, NJ (US);

Yu Zhang, Mountain View, CA (US);

Bhuvana Ramabhadran, Mt. Kisco, NY (US);

Pedro J. Moreno Mengibar, Jersey City, NJ (US);

Assignee:

Google LLC, Mountain View, CA (US);

Attorneys:
Primary Examiner:
Int. Cl.
CPC ...
G10L 15/00 (2013.01); G06F 40/289 (2020.01); G10L 15/06 (2013.01); G10L 15/16 (2006.01); G10L 15/197 (2013.01);
U.S. Cl.
CPC ...
G10L 15/063 (2013.01); G06F 40/289 (2020.01); G10L 15/16 (2013.01); G10L 15/197 (2013.01); G10L 2015/0635 (2013.01);
Abstract

A method includes receiving training data that includes unspoken textual utterances, un-transcribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances. Each unspoken textual utterance is not paired with any corresponding spoken utterance of non-synthetic speech. Each un-transcribed non-synthetic speech utterance not paired with a corresponding transcription. Each transcribed non-synthetic speech utterance paired with a corresponding transcription. The method also includes generating a corresponding alignment output for each unspoken textual utterance of the received training data using an alignment model. The method also includes pre-training an audio encoder on the alignment outputs generated for corresponding to the unspoken textual utterances, the un-transcribed non-synthetic speech utterances, and the transcribed non-synthetic speech utterances to teach the audio encoder to jointly learn shared speech and text representations.


Find Patent Forward Citations

Loading…