The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Feb. 10, 2026

Filed:

Jun. 30, 2023
Applicant:

Amazon Technologies, Inc., Seattle, WA (US);

Inventors:

Bastian Schnell, Berlin, DE;

Sri Vishnu Kumar Karlapati, Cambridge, GB;

Alexis Pierre Jean-Baptiste Moinet, Cambridge, GB;

Panagiota Karanasou, Cambridge, GB;

Thomas Renaud Drugman, Carnieres, BE;

Syed Ammar Abbas, Cambridge, GB;

Ewa Magdalena Muszynska, Cambridge, GB;

Assignee:

Amazon Technologies, Inc., Seattle, WA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G10L 13/08 (2013.01);
U.S. Cl.
CPC ...
G10L 13/08 (2013.01);
Abstract

A speech-processing system may be configured to generate expressive synthesized speech. The system may include a prosody prediction model that generates a combination of durations and acoustic representations that may be based on the content of the text as well as additional context information. The model may be trained to predict a joint probability between linguistic representations (e.g., derived from text) and combined duration/acoustic representations (e.g., derived from audio). At inference, the model can process linguistic representations derived from text to predict combined duration/acoustic representations. In some implementations, the model may process additional information; for example, semantic embeddings output by a language model based on the text. In another example, the model may receive a speaker embedding representing voice characteristics of a particular speaker. A decoder may process the durations and acoustic representations output by the model to generate audio data representing the synthesized speech and representing expressive prosodic variation.


Find Patent Forward Citations

Loading…