The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Jun. 01, 2021

Filed:

Mar. 27, 2019
Applicant:

Sri International, Menlo Park, CA (US);

Inventors:

Diego Castan Lavilla, Mountain View, CA (US);

Harry Bratt, Mountain View, CA (US);

Mitchell Leigh McLaren, Alderley, AU;

Assignee:

SRI INTERNATIONAL, Menlo Park, CA (US);

Attorneys:
Primary Examiner:
Int. Cl.
CPC ...
G10L 17/00 (2013.01); G10L 15/08 (2006.01); G10L 15/04 (2013.01); G10L 17/18 (2013.01); G10L 15/00 (2013.01); G10L 15/16 (2006.01); G10L 17/06 (2013.01);
U.S. Cl.
CPC ...
G10L 15/08 (2013.01); G10L 15/005 (2013.01); G10L 15/04 (2013.01); G10L 15/16 (2013.01); G10L 17/06 (2013.01); G10L 17/18 (2013.01);
Abstract

In an embodiment, the disclosed technologies include automatically recognizing speech content of an audio stream that may contain multiple different classes of speech content, by receiving, by an audio capture device, an audio stream; outputting, by one or more classifiers, in response to an inputting to the one or more classifiers of digital data that has been extracted from the audio stream, score data; where a score of the score data indicates a likelihood that a particular time segment of the audio stream contains speech of a particular class; where the one or more classifiers use one or more machine-learned models that have been trained to recognize audio of one or more particular classes to determine the score data; using a sliding time window process, selecting particular scores from the score data; using the selected particular scores, determining and outputting one or more decisions as to whether one or more particular time segments of the audio stream contain speech of one or more particular classes; where the one or more decisions are outputted within a real-time time interval of the receipt of the audio stream; where the one or more decisions are used by downstream processing of the audio stream to control any one or more of the following: labeling the audio stream, segmenting the audio stream, diarizing the audio stream.


Find Patent Forward Citations

Loading…