The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.
The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.
Patent No.:
Date of Patent:
Jun. 02, 2026
Filed:
Dec. 31, 2021
Sumant Milind Hanumante, Aliso Viejo, CA (US);
Qing Jin, Palo Alto, CA (US);
Sergei Korolev, Playa Vista, CA (US);
Denys Makoviichuk, Playa Vista, CA (US);
Jian Ren, Highland Park, NJ (US);
Dhritiman Sagar, Marina Del Rey, CA (US);
Patrick Timothy Mcsweeney Simons, Redondo Beach, CA (US);
Sergey Tulyakov, Santa Monica, CA (US);
Yang Wen, San Jose, CA (US);
Richard Zhuang, San Diego, CA (US);
Sumant Milind Hanumante, Aliso Viejo, CA (US);
Qing Jin, Palo Alto, CA (US);
Sergei Korolev, Playa Vista, CA (US);
Denys Makoviichuk, Playa Vista, CA (US);
Jian Ren, Highland Park, NJ (US);
Dhritiman Sagar, Marina Del Rey, CA (US);
Patrick Timothy Mcsweeney Simons, Redondo Beach, CA (US);
Sergey Tulyakov, Santa Monica, CA (US);
Yang Wen, San Jose, CA (US);
Richard Zhuang, San Diego, CA (US);
Snap Inc., Santa Monica, CA (US);
Abstract
Techniques for training a neural network having a plurality of computational layers with associated weights and activations for computational layers in fixed-point formats include determining an optimal fractional length for weights and activations for the computational layers; training a learned clipping-level with fixed-point quantization using a PACT process for the computational layers; and quantizing on effective weights that fuses a weight of a convolution layer with a weight and running variance from a batch normalization layer. A fractional length for weights of the computational layers is determined from current values of weights using the determined optimal fractional length for the weights of the computational layers. A fixed-point activation between adjacent computational layers is related using PACT quantization of the clipping-level and an activation fractional length from a node in a following computational layer. The resulting fixed-point weights and activation values are stored as a compressed representation of the neural network.