Selected reading: Scaling experiments, fitted relations and training budgets.
Open the source ↗How to study this source
How do neural networks learn representations?
Identify the fitted relationship, measurement conditions and limits of extrapolation.
Ideas and questions
Read background definitions when a term blocks the argument. Then return to the source and reconstruct its claim in your own words.
Read alongside, read against
Attention Is All You Need
Vaswani et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
BERT
Devlin et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Language Models are Few-Shot Learners
Brown et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Training Compute-Optimal Large Language Models
Hoffmann et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Deep Learning and the Information Bottleneck Principle
Tishby · Zaslavsky. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.