The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Sep. 24, 2019

Filed:

Oct. 12, 2017
Applicant:

Google Llc, Mountain View, CA (US);

Inventors:

Pararth Shah, Mountain View, CA (US);

Larry Paul Heck, Los Altos, CA (US);

Dilek Hakkani-Tur, Los Altos, CA (US);

Assignee:

GOOGLE LLC, Mountain View, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G10L 15/00 (2013.01); G10L 15/16 (2006.01); G10L 15/30 (2013.01); G10L 17/22 (2013.01); H04L 12/58 (2006.01); G06F 16/9032 (2019.01);
U.S. Cl.
CPC ...
G10L 15/30 (2013.01); G06F 16/90332 (2019.01); G10L 15/005 (2013.01); G10L 15/16 (2013.01); G10L 17/22 (2013.01); H04L 51/02 (2013.01);
Abstract

Techniques are described related to turn-based reinforcement learning for dialog management. In various implementations, dialog states and corresponding responsive actions generated during a multi-turn human-to-computer dialog session may be obtained. A plurality of turn-level training instances may be generated, each including: a given dialog state of the plurality of dialog states at an outset of a given turn of the human-to-computer dialog session; and a given responsive action that was selected based on the given dialog state. One or more of the turn-level training instances may further include a turn-level feedback value that reflects on the given responsive action selected during the given turn. A reward value may be generated based on an outcome of the human-to-computer dialog session. The dialog management policy model may be trained based on turn-level feedback values of the turn-level training instance(s) and the reward value.


Find Patent Forward Citations

Loading…