The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Sep. 15, 2026

Filed:

Jan. 28, 2025
Applicant:

Huazhong University of Science and Technology, Wuhan City, CN;

Inventors:

Song Wu, Wuhan City, CN;

Hao Wu, Wuhan City, CN;

Zhuo Huang, Wuhan City, CN;

Hao Fan, Wuhan City, CN;

Yue Yu, Wuhan City, CN;

Hai Jin, Wuhan City, CN;

Attorneys:
Primary Examiner:
Int. Cl.
CPC ...
G06T 1/20 (2006.01); G06T 1/60 (2006.01);
U.S. Cl.
CPC ...
G06T 1/20 (2013.01); G06T 1/60 (2013.01);
Abstract

A GPU-sharing method and apparatus for serverless inference loads is provided, wherein the method involves intercepting and forwarding GPU API calls made by inference tasks to an API proxy process to manage and allocate GPU resources. With a CPU and multiple GPUs connected through a bus, the GPUs communicate with the CPU only through an API proxy for process management and resource allocation of the GPUs. The CPU intercepts all GPU APIs triggered by any function of a same inference application, forwards the intercepted GPU APIs to a same designated GPU runtime for execution, and directs the GPU APIs triggered by each function to a pre-designated stream pool for the same inference application, so that all the functions of the same inference application share the same GPU runtime. The present disclosure solves the problem related to bulkiness of GPU runtimes in serverless inference systems, thereby facilitating GPU resource usage.


Find Patent Forward Citations

Loading…