Efficient Estimation of Word Representations in…

The Word2Vec paper is still one of the best first reads for representation learning. It shows how simple prediction tasks can turn words into vectors with useful geometric structure.

Reading focus: Why embeddings can encode similarity through learned context prediction. How CBOW and skip-gram frame word learning as efficient objectives. Why representation quality can matter more than model complexity.

ICLR Workshop 2013. Mikolov et al.. 30 min read, very easy difficulty.