The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Mar. 24, 2026

Filed:

May. 30, 2025
Applicant:

Intuit Inc., Mountain View, CA (US);

Inventors:

Ankita Sinha, Mountain View, CA (US);

Jiaxin Zhang, Mountain View, CA (US);

Kamalika Das, Saratoga, CA (US);

Sricharan Kallur Palli Kumar, Mountain View, CA (US);

Wendi Cui, Jersey City, NJ (US);

Assignee:

Intuit Inc., Mountain View, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06F 16/332 (2025.01); G06F 16/3329 (2025.01); G06F 16/338 (2019.01);
U.S. Cl.
CPC ...
G06F 16/3326 (2019.01); G06F 16/3329 (2019.01); G06F 16/338 (2019.01);
Abstract

A large language model (LLM) is refined using an unsupervised alignment process for enhanced reasoning and deep thinking capabilities. An iterative search and optimization procedure is used to explore the space of possible thought generations enabling the model to improve deep thinking capabilities without direct supervision. The LLM is prompted to produce a first thought and response pair in response to a query. The thought and response pair is evaluated with a trained judge LLM, which provides feedback for the thought and for the response. A revised prompt is generated in response to the feedback and provided to the LLM, which produces second thought and response pair. The LLM is refined based on the first, non-preferred, thought and response pair and the second, preferred, thought and response pair to align with the preferred thought and response pair, e.g., using a direct preference optimization process.


Find Patent Forward Citations

Loading…