The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Apr. 14, 2026

Filed:

May. 31, 2023
Applicant:

Google Llc, Mountain View, CA (US);

Inventors:

Harshit Kharbanda, Pleasanton, CA (US);

Belinda Luna Zeng, Cupertino, CA (US);

Viviana Caso Corella, San Francisco, CA (US);

Aashi Jain, Sunnyvale, CA (US);

David William Hendon, Oakland, CA (US);

Christopher James Kelley, Orinda, CA (US);

Jessica Lee, Brooklyn, NY (US);

Dounia Berrada, Saratoga, CA (US);

Kai Yu, San Francisco, CA (US);

Louis Wang, San Francisco, CA (US);

Thomas J. Duerig, Mountain View, CA (US);

Radu Soricut, Manhattan Beach, CA (US);

Robin Dua, San Francisco, CA (US);

Assignee:

GOOGLE LLC, Mountain View, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06T 7/70 (2017.01); G06F 16/732 (2019.01); G06F 16/735 (2019.01); G06F 16/783 (2019.01); G06V 10/62 (2022.01); G06V 10/774 (2022.01); G06V 20/40 (2022.01); G06V 10/82 (2022.01); G10L 15/26 (2006.01);
U.S. Cl.
CPC ...
G06F 16/735 (2019.01); G06F 16/732 (2019.01); G06F 16/7834 (2019.01); G06T 7/70 (2017.01); G06V 10/62 (2022.01); G06V 10/774 (2022.01); G06V 20/46 (2022.01); G06T 2207/10016 (2013.01); G06T 2207/20081 (2013.01); G06T 2207/20084 (2013.01); G06T 2207/30244 (2013.01); G06V 10/82 (2022.01); G06V 2201/07 (2022.01); G10L 15/26 (2013.01);
Abstract

A multimodal search system using a video query is described. The system can receive video data captured by a camera of a user device. The video data can have a sequence of image frames. Additionally, the system can receive audio data associated with the video data captured by the user device. Moreover, the system can process, using one or more machine-learned models, the sequence of image frames to generate video embeddings related to the sequence of the image frames. The video embeddings can have a plurality of image embeddings associated with the sequence of image frames. Furthermore, the system can determine one or more video results based on the video embeddings and the audio data. Subsequently, the system can transmit, to the user device, the one or more video results.


Find Patent Forward Citations

Loading…