Paper Discussions
In this section I practice actually reading academic papers — not just the math, but the claim, the evidence, and where I’d push back.
Papers
Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau et al., 2014) The paper that names the seq2seq bottleneck problem I found in my own translation experiment, and introduces attention as the fix.
Effective Approaches to Attention-based Neural Machine Translation (Luong et al., 2015) Following the Bahdanau et al. (2014) paper, this paper proposes a new attention mechanism called “global attention” and “local attention” that improves the performance of neural machine translation models.
Attention Is All You Need (Vaswani et al., 2017) The paper that introduces the Transformer architecture, which relies entirely on self-attention mechanisms and eliminates the need for recurrent or convolutional layers, leading to significant improvements in parallelization and performance for sequence-to-sequence tasks.
Language Models are Unsupervised Multitask Learners (Radford et al., 2019) The paper that introduces GPT-2, a large-scale unsupervised language model that demonstrates strong performance on a variety of natural language processing tasks, without task-specific training data, showcasing the potential of transfer learning in NLP.
Comments