Selected reading: Convolutional structure and attention mechanisms.
Open the source ↗How to study this source
How do neural networks learn representations?
Trace the attention mechanism and compare its structural and computational assumptions with convolution.
Ideas and questions
Read background definitions when a term blocks the argument. Then return to the source and reconstruct its claim in your own words.
Read alongside, read against
BERT
Devlin et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Language Models are Few-Shot Learners
Brown et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Scaling Laws for Neural Language Models
Kaplan et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Training Compute-Optimal Large Language Models
Hoffmann et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Deep Learning and the Information Bottleneck Principle
Tishby · Zaslavsky. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.