The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Mar. 24, 2026

Filed:

May. 01, 2023
Applicant:

Nvidia Corporation, Santa Clara, CA (US);

Inventors:

Jiarui Xu, San Diego, CA (US);

Shalini De Mello, San Francisco, CA (US);

Sifei Liu, Santa Clara, CA (US);

Arash Vahdat, Mountain View, CA (US);

Wonmin Byeon, Santa Cruz, CA (US);

Assignee:

NVIDIA Corporation, Santa Clara, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06V 20/70 (2022.01); G06T 7/10 (2017.01); G06V 10/40 (2022.01);
U.S. Cl.
CPC ...
G06T 7/10 (2017.01); G06V 10/40 (2022.01); G06T 2207/20081 (2013.01); G06T 2207/20084 (2013.01);
Abstract

An open-vocabulary diffusion-based panoptic segmentation system is not limited to perform segmentation using only object categories seen during training, and instead can also successfully perform segmentation of object categories not seen during training and only seen during testing and inferencing. In contrast with conventional techniques, a text-conditioned diffusion (generative) model is used to perform the segmentation. The text-conditioned diffusion model is pre-trained to generate images from text captions, including computing internal representations that provide spatially well-differentiated object features. The internal representations computed within the diffusion model comprise object masks and a semantic visual representation of the object. The semantic visual representation may be extracted from the diffusion model and used in conjunction with a text representation of a category label to classify the object. Objects are classified by associating the text representations of category labels with the object masks and their semantic visual representations to produce panoptic segmentation data.


Find Patent Forward Citations

Loading…