Efficient Estimation of Word Representations in Vector Space
We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks.
word2vec has no earlier papers in this dataset.
Led to
- Weakly supervised visual features (Jouli2015 · cited 6דWe follow Mikolov et al. mikolov13 and sample instances uniformly per class.”From Weakly supervised visual features (Jouli · §Weakly Supervised Learning of Convnets
- CPC2018 · cited 2דRecent work in unsupervised learning has successfully used these ideas to learn word representations by predicting neighboring words [9].”From CPC · §Introduction
- Gopher2021 · cited 1×, 1 in Method“Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”From Gopher · §Background
- GLaM2021 · cited 2דWord embedding models and extensions such as word2vec (Mikolov et al. 2013), GloVe (Pennington et al. 2014) and paragraph vectors (Le & Mikolov 2014) have shown good generalization to many tasks simply by transferring th…”From GLaM · §Related Work
Abstract
We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.