Open Access
Measuring Brazilian Portuguese Product Titles Similarity using Embeddings
Peer ReviewedAlan da Silva Romualdo +22021Conference proceedings
Textual similarity deals with determining how similar two pieces of texts are, considering the lexical (surface forms) or semantic (meaning) closeness. In this paper we applied word embeddings for measuring e-commerce product title similarity in Brazilian Portuguese. We generated some domainspecific word embeddings (using Word2Vec, FastText and GloVe) and compared them with general-domain models (word embeddings and BERT models). We concluded that the cosine similarity calculated using the domain-specific word embeddings was a good approach to distinguish between similar and nonsimilar products, but the multilingual BERT pre-trained model proved to be the best one.

The content you want is available to Zendy users.

Already have an account? Sign in
Having issues? Contact support