Selected reading: Full paper; methods and limitations.
Open the source ↗How to study this source
How do neural networks learn representations?
Compare this account’s mechanism with the preceding reading; note where their predictions differ.
Ideas and questions
Read background definitions when a term blocks the argument. Then return to the source and reconstruct its claim in your own words.
Read alongside, read against
Attention Is All You Need
Vaswani et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Language Models are Few-Shot Learners
Brown et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Scaling Laws for Neural Language Models
Kaplan et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Training Compute-Optimal Large Language Models
Hoffmann et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Deep Learning and the Information Bottleneck Principle
Tishby · Zaslavsky. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.