The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.

The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.

Date of Patent:
Jun. 23, 2026

Filed:

Jun. 19, 2023
Applicant:

Google Llc, Mountain View, CA (US);

Inventors:

Seungyeon Kim, New York, NY (US);

Ankit Singh Rawat, Jersey City, NJ (US);

Wittawat Jitkrittum, Jersey City, NJ (US);

Hari Narasimhan, Mountain View, CA (US);

Sashank Reddi, Fort Lee, NJ (US);

Neha Gupta, New York, NY (US);

Srinadh Bhojanapalli, Maplewoood, NJ (US);

Aditya Menon, New York, NY (US);

Manzil Zaheer, Mountain View, CA (US);

Tal Schuster, New York, NY (US);

Sanjiv Kumar, Jericho, NY (US);

Toby Boyd, Columbus, OH (US);

Zhifeng Chen, Sunnyvale, CA (US);

Emanuel Taropa, Los Altos, CA (US);

Vikram Kasivajhula, San Francisco, CA (US);

Trevor Strohman, Sunnyvale, CA (US);

Martin Baeuml, Zurich, CH;

Leif Schelin, Zurich, CH;

Yanping Huang, Mountain View, CA (US);

Assignee:

GOOGLE LLC, Mountain View, CA (US);

Attorney:
Primary Examiner:
Int. Cl.
CPC ...
G06F 16/3329 (2025.01);
U.S. Cl.
CPC ...
G06F 16/3329 (2019.01);
Abstract

Implementations disclose selecting, in response to receiving a request and from among multiple candidate generative models (e.g., multiple candidate large language models (LLMs)) with differing computational efficiencies, a particular generative model to utilize in generating a response to the request. Those implementations reduce latency and/or conserve computational resource(s) through selection, for various requests, of a more computationally efficient generative model for utilization in lieu of a less computationally efficient generative model. Further, those implementations seek to achieve such benefits, through utilization of more computationally efficient generative models, while also still selectively utilizing less computationally efficient generative models for certain requests to mitigate occurrences of a generated response being inaccurate and/or under-specified. This, in turn, can mitigate occurrences of computational and/or network inefficiencies that result from a user issuing a follow-up request to cure the inaccuracies and/or under-specification of a generated response.


Find Patent Forward Citations

Loading…