The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Sep. 22, 2026

Filed:

Apr. 25, 2024
Applicant:

Google Llc, Mountain View, CA (US);

Inventors:

Quan Wang, Hoboken, NJ (US);

Yiling Huang, Edgewater, NJ (US);

Guanlong Zhao, Long Island City, NY (US);

Assignee:

Google LLC, Mountain View, CA (US);

Attorneys:
Primary Examiner:
Int. Cl.
CPC ...
G06F 40/284 (2020.01); G10L 15/06 (2013.01); G10L 15/07 (2013.01); G10L 17/02 (2013.01);
U.S. Cl.
CPC ...
G06F 40/284 (2020.01); G10L 15/063 (2013.01); G10L 15/07 (2013.01); G10L 17/02 (2013.01);
Abstract

A method includes receiving a prompt including a textual diarization request and corresponding audio data characterizing a conversation between multiple speakers. The method also includes generating a sequence of audio encoding chunks based on the corresponding data. For each respective audio encoding chunk, the method includes using a trained large language model (LLM) generating corresponding diarization results based on the respective audio encoding chunk and the textual diarization request and generating a new audio cohort for the respective audio encoding chunk based on the corresponding diarization results. The corresponding diarization results include a speech recognition result that has one or more predicted terms. Each respective predicted term is associated with a corresponding speaker token representing a predicted identity of a respective speaker that spoke the respective predicted term. The trained LLM is conditioned on a prior audio cohort generated by the trained LLM for a prior audio encoding chunk.


Find Patent Forward Citations

Loading…