The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Jun. 30, 2026

Filed:

Jul. 28, 2025
Applicant:

Intuit Inc., Mountain View, CA (US);

Inventors:

Sagiv Antebi, Tel Aviv, IL;

Matan Vetzler, Tel Aviv, IL;

Shai Ardazi, Tel Aviv, IL;

Ofir Ben Shoham, Tel Aviv, IL;

Assignee:

Intuit Inc., Mountain View, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06N 3/0895 (2023.01); G06N 3/02 (2006.01); G06N 3/08 (2023.01); G06N 3/092 (2023.01); G06N 3/094 (2023.01); G06N 20/00 (2019.01);
U.S. Cl.
CPC ...
G06N 3/0895 (2023.01); G06N 3/02 (2013.01); G06N 3/08 (2013.01); G06N 3/092 (2023.01); G06N 3/094 (2023.01); G06N 20/00 (2019.01);
Abstract

Certain aspects of the disclosure provide a method for training a language model (LM) including: generating, using an LM, one or more outputs; computing a confidence score of an output of the one or more outputs based on a perplexity value of the output; determining, by a group relative policy optimization (GRPO)-based model, that the output is: associated with a correct status based on a reference policy; and associated with an uncertain status based on the confidence score and a threshold; determining, by the GRPO-based model, an increased reward value for the output that is associated with the correct status and the uncertain status based at least in part on a base reward value and the confidence score; causing, by the GRPO-based model, a reinforcement of the output using the increased reward value; and training the LM in accordance with the reinforcement of the output.


Find Patent Forward Citations

Loading…