The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Jan. 31, 2023

Filed:

Apr. 28, 2019
Applicant:

International Business Machines Corporation, Armonk, NY (US);

Inventors:

Eitan Farchi, Pardes Hanna, IL;

Orna Raz, Haifa, IL;

Marcel Zalmanovici, Kiriat Motzkin, IL;

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06F 17/00 (2019.01); G06K 9/62 (2022.01); G06N 20/00 (2019.01);
U.S. Cl.
CPC ...
G06K 9/6256 (2013.01); G06K 9/6218 (2013.01); G06N 20/00 (2019.01);
Abstract

A method, apparatus and product for identifying data drifts. The method comprising: obtaining a baseline dataset of instances in a feature space, each of wherein being associated with a label; determining a set of clusters in the feature space, based on the baseline dataset; determining a baseline distribution of instances over the set of clusters based on the baseline dataset; for each cluster, computing a performance metric for a predictor for the each cluster, wherein the predictor is configured to estimate an estimated label for an instance, wherein the performance metric is indicative of a successful estimation of the predictor to a portion of the baseline dataset that is comprised by the cluster; obtaining a second dataset, wherein the second dataset comprising instances in the feature space; determining a second distribution of instances over the set of clusters, wherein said determining the second distribution is based on the second dataset; and based on the second distribution and on the baseline distribution, and based on at least one performance metric of at least one the cluster of the set of clusters, identifying a data drift in the second dataset with respect to the baseline dataset.


Find Patent Forward Citations

Loading…