The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.
The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.
Patent No.:
Date of Patent:
May. 12, 2026
Filed:
Apr. 23, 2024
Meta-reflection techniques for learning instructions for language agents using past self-reflections
Microsoft Technology Licensing, Llc, Redmond, WA (US);
Gustavo Araujo Soares, Redmond, WA (US);
Sumit Gulwani, Sammamish, WA (US);
Shashank Kirtania, Bengaluru, IN;
Sherry Shi, Redmond, WA (US);
Arjun Radhakrishna, Seattle, WA (US);
Ananya Singha, Ranchi, IN;
Priyanshu Gupta, Bengaluru, IN;
Microsoft Technology Licensing, LLC, Redmond, WA (US);
Abstract
A data processing system implements accessing a datastore of training data using a model training unit to obtain a first training sample, the first training sample comprising a first natural language utterance, first ground truth information, the first natural language utterance requesting that content be generated by a language model, the first ground truth information providing a first example of first expected output of the language model in response to the first natural language utterance; constructing a first prompt based on the first natural language utterance using a prompt construction unit; providing, using the prompt construction unit, the first prompt to the language model as an input to cause the language model to generate a first output; analyzing the first output and the first ground truth information using the model training unit to determine whether the first output is erroneous; constructing, using the prompt construction unit, a second prompt that instructs the language model to generate a first self-reflection response that indicates why the language model generated the first output; providing the second prompt as an input to the language model to cause the language model to generate the first self-reflection response; constructing, using the prompt construction unit, a third prompt that includes the first self-reflection response, the third prompt instructing the language model to generate prompt improvement instructions to be included in subsequently constructed prompts for the language model to assist the language model in generating a correct response to the subsequently constructed prompts; providing the third prompt to the language model to cause the language model to generate the prompt improvement instructions; and including the prompt improvement instructions in the subsequently constructed prompts generated using the prompt construction unit.