Gradient-Based Learning Applied to Document Recognition

LeCun et al.1998paper

Selected reading: Convolutional structure and attention mechanisms.

Open the source ↗

How to study this source

Ideas and questions

Read background definitions when a term blocks the argument. Then return to the source and reconstruct its claim in your own words.

Read alongside, read against

  • Attention Is All You Need

    Vaswani et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.

  • BERT

    Devlin et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.

  • Language Models are Few-Shot Learners

    Brown et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.

  • Scaling Laws for Neural Language Models

    Kaplan et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.

  • Training Compute-Optimal Large Language Models

    Hoffmann et al.. Compare assumptions, evidence and scope with the source above. These are editorial companions, not necessarily direct responses.