Reference synthesis
An editorial comparison of the route’s selected accounts and evidence.
A curated synthesis cannot be neutral or exhaustive. Follow the primary sources when an interpretation matters.
Finding the next step…
Build from differentiation and optimisation to architectures, generalisation and scaling, using small reproducible experiments.
Python, vectors, derivatives and elementary probability are prerequisites for the worked tasks. The Python, Linear algebra, Calculus and Probability routes provide preparation.
A trained model with a baseline, held-out evaluation and failure analysis.
Start here if the background is new. Equivalent experience is enough.
Use Python functions, arrays and tests in the model-training tasks.
Use vectors, matrix operations and geometric intuition to explain a model’s calculations.
Use derivatives and gradients to explain how a training update changes the model.
Use probability and uncertainty to examine model outputs and held-out evaluation.
Ian Goodfellow, Yoshua Bengio · Aaron Courville · reference
Task definitions, data splits and vector shapes
Check reading access
Open the readingIan Goodfellow, Yoshua Bengio · Aaron Courville · reference
Task definitions, data splits and vector shapes
Check reading access
Open the readingDeep Learning: Machine learning basics. Define the prediction task, baseline and evaluation split before choosing a model.
Deep Learning: Linear algebra. Track the shapes of inputs, weights and outputs in a tiny network.
Ian Goodfellow, Yoshua Bengio · Aaron Courville · reference
Feedforward computation, loss and chain-rule derivatives
Check reading access
Open the readingRumelhart, Hinton · Williams · paper · 1986
Feedforward computation, loss and chain-rule derivatives
Check reading access
Open the readingDeep Learning: Deep feedforward networks. Follow the forward computation and loss through one small network.
Learning Representations by Back-Propagating Errors. Trace how a change in one weight affects the loss through the chain rule.
Ian Goodfellow, Yoshua Bengio · Aaron Courville · reference
A small optimisation example and regularisation comparison
Check reading access
Open the readingIan Goodfellow, Yoshua Bengio · Aaron Courville · reference
A small optimisation example and regularisation comparison
Check reading access
Open the readingDeep Learning: Optimisation. Step through a small optimisation example and record the update assumptions.
Deep Learning: Regularisation. Compare training and held-out error before explaining what a regulariser changes.
Kaplan et al. · paper · 2020
Scaling experiments, fitted relations and training budgets
Check reading access
Open the readingHoffmann et al. · paper · 2022
Scaling experiments, fitted relations and training budgets
Check reading access
Open the readingScaling Laws for Neural Language Models. Identify the fitted relationship, measurement conditions and limits of extrapolation.
Training Compute-Optimal Large Language Models. Compare model size, training data and compute assumptions with the earlier scaling account.
Editorial perspectives based on selected works, rather than author-endorsed reading lists.
An editorial comparison of the route’s selected accounts and evidence.
A curated synthesis cannot be neutral or exhaustive. Follow the primary sources when an interpretation matters.
An editorial reconstruction using the selected work of Geoffrey Hinton.
This is not an author-endorsed syllabus. The reconstruction highlights selected works and may omit other commitments.
An editorial reconstruction using the selected work of Yann LeCun.
This is not an author-endorsed syllabus. The reconstruction highlights selected works and may omit other commitments.
An editorial reconstruction using the selected work of Ilya Sutskever.
This is not an author-endorsed syllabus. The reconstruction highlights selected works and may omit other commitments.
Background, different viewpoints, and further reading.
Trace how a change in one weight affects the loss through the chain rule.
Feedforward computation, loss and chain-rule derivatives
Compare this account’s mechanism with the preceding reading; note where their predictions differ.
Selected chapters and argument
Identify the input structure, learned features and benchmark comparison in the selected example.
Convolutional structure and attention mechanisms
Use this account to revise your initial explanation and identify an unresolved question.
Full paper; methods and limitations
Trace the attention mechanism and compare its structural and computational assumptions with convolution.
Convolutional structure and attention mechanisms
Compare this account’s mechanism with the preceding reading; note where their predictions differ.
Full paper; methods and limitations
Extract one claim and distinguish the evidence supporting it from the author’s interpretation.
Full paper; methods and limitations
Identify the fitted relationship, measurement conditions and limits of extrapolation.
Scaling experiments, fitted relations and training budgets
Compare model size, training data and compute assumptions with the earlier scaling account.
Scaling experiments, fitted relations and training budgets
Compare this account’s mechanism with the preceding reading; note where their predictions differ.
Full paper; methods and limitations
Extract one claim and distinguish the evidence supporting it from the author’s interpretation.
Full paper; methods and limitations
Use this account to revise your initial explanation and identify an unresolved question.
Selected chapters and argument
Track the shapes of inputs, weights and outputs in a tiny network.
Task definitions, data splits and vector shapes
Compare entropy, cross-entropy and KL divergence.
Named section and worked examples
Inspect stability and conditioning.
Named section and worked examples
Define the prediction task, baseline and evaluation split before choosing a model.
Task definitions, data splits and vector shapes
Follow the forward computation and loss through one small network.
Feedforward computation, loss and chain-rule derivatives
Compare training and held-out error before explaining what a regulariser changes.
A small optimisation example and regularisation comparison
Step through a small optimisation example and record the update assumptions.
A small optimisation example and regularisation comparison
Trace receptive fields and equivariance.
Named section and worked examples
Contrast held-out predictive performance with the validity of an intervention claim.