The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.
The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.
Patent No.:
Date of Patent:
Jun. 30, 2026
Filed:
Oct. 17, 2025
Openai Opco, Llc, San Francisco, CA (US);
John Allard, San Francisco, CA (US);
James Blomo, San Carlos, CA (US);
Filipe De Avila Belbute Peres, San Francisco, CA (US);
William Hang, San Francisco, CA (US);
Joseph Palermo, San Francisco, CA (US);
Chaitanya Ravuri, Saratoga, CA (US);
Theophile Sautory, San Francisco, CA (US);
Karan Sharma, San Francisco, CA (US);
Beining Zhou, Stanford, CA (US);
Wenjie Zi, San Francisco, CA (US);
OpenAI OpCo, LLC, San Francisco, CA (US);
Abstract
The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing a reinforcement fine-tuning pipeline to fine tune an agent machine learning (ML) model utilizing user-defined grader endpoints and/or external tools with a multi-layer security architecture. For example, the disclosed systems can utilize a multi-layer security architecture to filter incoming training datasets, perform content refusal checks, chain-of-thought leak detections, and/or governance oversights prior to training the agent ML model, during active training of the agent ML model using external tools and/or grader models, and/or during post-training of the agent ML model (prior to releasing a fine-tuned snapshot of the model). In addition, the disclosed systems can generate stateful trajectory rollouts to associate training trajectories of the agent machine learning model to unique identifiers to facilitate multiple environments and/or tasks to run concurrently while maintaining consistency across external tool calls and grading interactions during training.