
Sander Dieleman, Research Scientist (Director) at Google DeepMind, walks through a short history of depth in deep learning. Each breakthrough let networks go deeper, from two layers to five, then to 20, then to 100. Sander sees diffusion and autoregression as the next step in that same line of progress. In this clip, Sander covers: • How layer-wise pre-training took networks from two layers to five • Why ReLU and other non-saturating nonlinearities pushed depth to 10 to 20 layers • How residual connections made 100-layer networks possible • Why he frames diffusion and autoregression as the next step in that evolution





