Definition
with the random initialization : how similarly the predictions at and change under a small parameter update.
Formula
Gradient flow with squared loss: , . Very wide: .
In the NTK regime the weights barely move, the tangent features stay fixed and the network behaves like a random feature model or kernel method. NTK theory explains optimization of very wide networks, not feature learning in practical ones.