The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Oct. 26, 2021

Filed:

Mar. 31, 2021
Applicant:

Alipay (Hangzhou) Information Technology Co., Ltd., Zhejiang, CN;

Inventors:

Hui Li, Hangzhou, CN;

Le Song, Hangzhou, CN;

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06F 9/46 (2006.01); G06F 9/48 (2006.01); G06N 3/08 (2006.01); G06K 9/62 (2006.01); B60W 60/00 (2020.01); G06F 9/52 (2006.01); G06N 3/02 (2006.01);
U.S. Cl.
CPC ...
G06F 9/4881 (2013.01); B60W 60/00 (2020.02); G06F 9/4843 (2013.01); G06F 9/528 (2013.01); G06K 9/6257 (2013.01); G06K 9/6277 (2013.01); G06N 3/02 (2013.01); G06N 3/08 (2013.01); G05B 2219/25255 (2013.01);
Abstract

Computer-implemented methods, systems, and apparatus, including computer-readable medium, for generating an action selection policy for causing an execution device to complete a task are described. Data representing a task that is divided into a sequence of subtasks are obtained. For a specified subtask except for a first subtask in the sequence of subtasks, a value neural network (VNN) is trained. The VNN receives inputs include reach probabilities of reaching a subtask initial state of the specified subtask, and predicts a reward of the execution device in the subtask initial state of the specified subtask. A strategy neural network (SNN) for a prior subtask that precedes the specified subtask is trained based on the VNN. The SNN receives inputs include a sequence of actions that reach a subtask state of the prior subtask, and predicts an action selection policy of the execution device in the subtask state of the prior subtask.


Find Patent Forward Citations

Loading…