Effective Theory of Deep Neural Networks

20 Apr 2022

Effective Theory of Deep Neural Networks

Sho Yaida, Meta AI

Abstract:
Large neural networks perform extremely well in practice, providing the backbone of modern machine learning. The goal of this talk is to provide a blueprint for theoretically analyzing these large models from first principles. In particular, we’ll overview how the statistics and dynamics of deep neural networks drastically simplify at large width and become analytically tractable. In so doing, we’ll see that the idealized infinite-width limit is too simple to capture several important aspects of deep learning such as representation learning. To address them, we’ll step beyond the idealized limit and systematically incorporate finite-width corrections.

video

Series

seminars

Physics ∩ ML

Effective Theory of Deep Neural Networks