The patent badge is an abbreviated version of the USPTO patent document. The patent badge does contain a link to the full patent document.
The patent badge is an abbreviated version of the USPTO patent document. The patent badge covers the following: Patent number, Date patent was issued, Date patent was filed, Title of the patent, Applicant, Inventor, Assignee, Attorney firm, Primary examiner, Assistant examiner, CPCs, and Abstract. The patent badge does contain a link to the full patent document (in Adobe Acrobat format, aka pdf). To download or print any patent click here.
Patent No.:
Date of Patent:
Apr. 14, 2026
Filed:
Nov. 04, 2022
Oracle International Corporation, Redwood Shores, CA (US);
Thanh Tien Vu, Herston, AU;
Poorya Zaremoodi, Melbourne, AU;
Duy Vu, Melbourne, AU;
Mark Edward Johnson, Sydney, AU;
Thanh Long Duong, Seabrook, AU;
Xu Zhong, Melbourne, AU;
Vladislav Blinov, Melbourne, AU;
Cong Duy Vu Hoang, Wantirna South, AU;
Yu-Heng Hong, Carlton, AU;
Vinamr Goel, Ghaziabad, IN;
Philip Victor Ogren, Boulder, CO (US);
Srinivasa Phani Kumar Gadde, Fremont, CA (US);
Vishal Vishnoi, Redwood City, CA (US);
Oracle International Corporation, Redwood Shores, CA (US);
Abstract
Techniques disclosed herein relate generally to language detection. In one particular aspect, a method is provided that includes obtaining a sequence of n-grams of a textual unit; using an embedding layer to obtain an ordered plurality of embedding vectors for the sequence of n-grams; using a deep network to obtain an encoded vector that is based on the ordered plurality of embedding vectors; and using a classifier to obtain a language prediction for the textual unit that is based on the encoded vector. The deep network includes an attention mechanism, and using the embedding layer to obtain the ordered plurality of embedding vectors comprises, for each n-gram in the sequence of n-grams: obtaining hash values for the n-gram; based on the hash values, selecting component vectors from among the plurality of component vectors; and obtaining an embedding vector for the n-gram that is based on the component vectors.