The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.
The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.
Patent No.:
Date of Patent:
Jul. 28, 2026
Filed:
Jan. 29, 2025
Oracle International Corporation, Redwood Shores, CA (US);
Thanh Tien Vu, Herston, AU;
Prashan Madumal, Melbourne, AU;
Omid Mohamad Nezami, Sydney, AU;
Cong Duy Vu Hoang, Wantirna South, AU;
Gioacchino Tangari, Sydney, AU;
Duy Vu, Melbourne, AU;
Dai Quoc Nguyen, Raceview, AU;
Yakupitiyage Don Thanuja Samodhye Dharmasiri, Melbourne, AU;
Ying Xu, Melbourne, AU;
Thanh Long Duong, Melbourne, AU;
Clemence Andree Simone Lanfranchi, Zurich, CH;
Damien Alexandre Hilloulin, Zurich, CH;
Rhicheek Patra, Zurich, CH;
Sungpack Hong, Palo Alto, CA (US);
Hassan Chafi, Zurich, CH;
Oracle International Corporation, Redwood Shores, CA;
Abstract
Techniques are disclosed herein for automatically curating preference alignment training (PAT) data and subsequently using the PAT data to train a machine-learning model used to perform natural language-to-query language tasks. A first PAT dataset is generated using a first trained LLM in an execution-based data generation technique. A second PAT dataset is generated using a second trained LLM in a LLM-based data generation technique. The second PAT dataset is populated with synthetic data points generated by the second trained LLM. Each data point of first PAT dataset and the second PAT dataset comprises a natural language query and an executable and non-executable query language statement. The first PAT dataset and the second PAT dataset can be combined and filtered to generate a PAT dataset that is optimized for training a machine-learning model to overcome one or more weaknesses exhibited when generating query language statements responsive to natural language questions.