The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Sep. 15, 2026

Filed:

Feb. 22, 2024
Applicant:

Google Llc, Mountain View, CA (US);

Inventors:

Wenqian Huang, Mountain View, CA (US);

Hao Zhang, Jericho, NY (US);

Shankar Kumar, New York, NY (US);

Shuo-Yiin Chang, Sunnyvale, CA (US);

Tara N. Sainath, Jersey City, NJ (US);

Assignee:

Google LLC, Mountain View, CA (US);

Attorneys:
Primary Examiner:
Int. Cl.
CPC ...
G10L 15/16 (2006.01); G06F 40/30 (2020.01); G10L 15/06 (2013.01); G10L 15/22 (2006.01); G10L 15/26 (2006.01);
U.S. Cl.
CPC ...
G10L 15/063 (2013.01); G06F 40/30 (2020.01); G10L 15/26 (2013.01);
Abstract

A joint segmenting and ASR model includes an encoder to receive a sequence of acoustic frames and generate, at each of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame. The model also includes a decoder to generate based on the higher order feature representation at each of the plurality of output steps a probability distribution over possible speech recognition hypotheses, and an indication of whether the corresponding output step corresponds to an end of segment (EOS). The model is trained on a set of training samples, each training sample including audio data characterizing multiple segments of long-form speech; and a corresponding transcription of the long-form speech, the corresponding transcription annotated with ground-truth EOS labels obtained via distillation from a language model teacher that receives the corresponding transcription as input and injects the ground-truth EOS labels into the corresponding transcription between semantically complete segments.


Find Patent Forward Citations

Loading…