The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Oct. 08, 2024

Filed:

Jan. 31, 2022
Applicant:

Salesforce, Inc., San Francisco, CA (US);

Inventors:

Shu Zhang, Fremont, CA (US);

Junnan Li, Singapore, SG;

Ran Xu, Mountain View, CA (US);

Caiming Xiong, Menlo Park, CA (US);

Chetan Ramaiah, San Bruno, CA (US);

Assignee:

Salesforce, Inc., San Francisco, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06V 10/776 (2022.01); G06F 16/56 (2019.01); G06F 16/583 (2019.01); G06F 40/126 (2020.01); G06F 40/166 (2020.01); G06F 40/284 (2020.01); G06V 10/74 (2022.01); G06V 10/80 (2022.01);
U.S. Cl.
CPC ...
G06V 10/776 (2022.01); G06F 16/56 (2019.01); G06F 16/5846 (2019.01); G06F 40/126 (2020.01); G06F 40/166 (2020.01); G06F 40/284 (2020.01); G06V 10/761 (2022.01); G06V 10/806 (2022.01);
Abstract

Embodiments described herein a CROss-Modal Distribution Alignment (CROMDA) model for vision-language pretraining, which can be used for retrieval downstream tasks. In the CROMDA mode, global cross-modal representations are aligned on each unimodality. Specifically, a uni-modal global similarity between an image/text and the image/text feature queue are computed. A softmax-normalized distribution is then generated based on the computed similarity. The distribution thus takes advantage of property of the global structure of the queue. CROMDA then aligns the two distributions and learns a modal invariant global representation. In this way, CROMDA is able to obtain invariant property in each modality, where images with similar text representations should be similar and vice versa.


Find Patent Forward Citations

Loading…