The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Dec. 16, 2025

Filed:

Nov. 09, 2022
Applicant:

Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, AE;

Inventors:

Omkar Thawakar, Abu Dhabi, AE;

Sanath Narayan, Abu Dhabi, AE;

Hisham Cholakkal, Abu Dhabi, AE;

Rao Muhammad Anwer, Abu Dhabi, AE;

Muhammad Haris, Abu Dhabi, AE;

Salman Khan, Abu Dhabi, AE;

Fahad Khan, Abu Dhabi, AE;

Attorney:
Primary Examiner:
Assistant Examiner:
Int. Cl.
CPC ...
G06T 7/73 (2017.01); G06V 10/26 (2022.01); G06V 10/77 (2022.01); G06V 10/774 (2022.01); G06V 10/94 (2022.01); G06V 20/58 (2022.01);
U.S. Cl.
CPC ...
G06T 7/73 (2017.01); G06V 10/26 (2022.01); G06V 10/7715 (2022.01); G06V 10/774 (2022.01); G06V 10/95 (2022.01); G06V 20/582 (2022.01); G06T 2207/30196 (2013.01); G06T 2207/30261 (2013.01);
Abstract

A system, method, computer readable storage medium for a computer vision system includes at least one video camera, and video processor circuitry. The method includes inputting a stream of video data and generating a sequence of image frames, segmenting and tracking, by the video analysis apparatus, object instances in the stream of video data, including receiving the sequence of image frames, analyzing the sequence of image frames using a video instance segmentation transformer to obtain a video instance mask sequence from the sequence of image frames, the transformer having a backbone network, a transformer encoder-decoder, and an instance matching and segmentation block, The encoder contains a multi-scale spatio-temporal split attention module to capture spatio-temporal feature relationships at multiple scales across multiple frames. The decoder contains a temporal attention block for enhancing a temporal consistency of transformer queries. The method includes displaying the video instance mask sequence.


Find Patent Forward Citations

Loading…