The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Sep. 08, 2026

Filed:

Feb. 04, 2022
Applicant:

Adobe Inc., San Jose, CA (US);

Inventors:

Yaman Kumar, New Delhi, IN;

Balaji Krishnamurthy, Noida, IN;

Assignee:

Adobe Inc., San Jose, CA (US);

Attorney:
Primary Examiner:
Assistant Examiner:
Int. Cl.
CPC ...
G10L 15/25 (2013.01); G06F 18/23213 (2023.01); G06N 3/02 (2006.01); G06T 9/00 (2006.01); G06V 10/82 (2022.01); G06V 20/40 (2022.01); G10L 13/02 (2013.01); G10L 15/16 (2006.01); G10L 15/22 (2006.01); G10L 25/57 (2013.01);
U.S. Cl.
CPC ...
G10L 15/25 (2013.01); G06F 18/23213 (2023.01); G06N 3/02 (2013.01); G06T 9/002 (2013.01); G06V 10/82 (2022.01); G06V 20/49 (2022.01); G10L 13/02 (2013.01); G10L 15/16 (2013.01); G10L 15/22 (2013.01); G10L 25/57 (2013.01);
Abstract

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that recognize speech from a digital video utilizing an unsupervised machine learning model, such as a generative adversarial neural network (GAN) model. In one or more implementations, the disclosed systems utilize an image encoder to generate self-supervised deep visual speech representations from frames of an unlabeled (or unannotated) digital video. Subsequently, in one or more embodiments, the disclosed systems generate viseme sequences from the deep visual speech representations (e.g., via segmented visemic speech representations from clusters of the deep visual speech representations) utilizing the adversarially trained GAN model. Indeed, in some instances, the disclosed systems decode the viseme sequences belonging to the digital video to generate an electronic transcription and/or digital audio for the digital video.


Find Patent Forward Citations

Loading…