The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Sep. 20, 2022

Filed:

Mar. 09, 2020
Applicant:

Kabushiki Kaisha Toshiba, Tokyo, JP;

Inventors:

Ryosuke Nonaka, Yokohama Kangawa, JP;

Toshimitsu Kaneko, Kawasaki Kanagawa, JP;

Norihiro Nakamura, Kawasaki Kanagawa, JP;

Assignee:
Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06N 20/00 (2019.01); B60W 30/18 (2012.01); B60W 30/188 (2012.01); G05D 1/02 (2020.01); B60W 50/00 (2006.01);
U.S. Cl.
CPC ...
G06N 20/00 (2019.01); B60W 30/188 (2013.01); B60W 30/18163 (2013.01); B60W 50/00 (2013.01); G05D 1/0221 (2013.01); B60W 2050/0075 (2013.01); G05D 2201/0213 (2013.01);
Abstract

According to one embodiment, a learning method, comprises receiving a first signal including a previous auxiliary variable value, previous action information regarding a previous action, or a set of previous scores, receiving current sensor data, selecting a current action of the control target based on the first signal, the current sensor data, and a parameter for obtaining a score from sensor data, causing the control target to execute the current action, receiving next sensor data and a reward, and updating the parameter based on the current sensor data, current action information regarding the current action, the next sensor data, and the reward. A degree of selecting a previous action as the current action is increased.


Find Patent Forward Citations

Loading…