Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory? - Citegraph

Paper Info

Title
Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory?

Abstract
Neural Tangent Kernel (NTK) theory is widely used to study the dynamics of infinitely-wide deep neural networks (DNNs) under gradient descent. But do the results for infinitely-wide networks give us hints about the behaviour of real finite-width ones? In this paper we study empirically when NTK theory is valid in practice for fully-connected ReLu and sigmoid networks. We find out that whether a network is in the NTK regime depends on the hyperparameters of random initialization and network's depth. In particular, NTK theory does not explain behaviour of sufficiently deep networks initialized so that their gradients explode: the kernel is random at initialization and changes significantly during training, contrary to NTK theory. On the other hand, in case of vanishing gradients DNNs are in the NTK regime but become untrainable rapidly with depth. We also describe a framework to study generalization properties of DNNs by means of NTK theory and discuss its limits.

Year	Venue	DocType
2021	Mathematical and Scientific Machine Learning (MSML)	Conference
Citations	PageRank	References
0	0.34	0
Authors
2

Authors (2 rows)

Cited by (0 rows)

References (0 rows)

Name	Order	Citations	PageRank
Mariia Seleznova	1	0	0.68
Gitta Kutyniok	2	325	34.77

1