Record learning-rate sensitivity.
Open the source ↗How to study this source
How do neural networks learn representations?
Step through a small optimisation example and record the update assumptions.
Ideas and questions
Read background definitions when a term blocks the argument. Then return to the source and reconstruct its claim in your own words.
Read alongside, read against
Attention Is All You Need
Vaswani et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
BERT
Devlin et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Language Models are Few-Shot Learners
Brown et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Scaling Laws for Neural Language Models
Kaplan et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.
Training Compute-Optimal Large Language Models
Hoffmann et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.