Definition

with the random initialization : how similarly the predictions at and change under a small parameter update.

Formula

Gradient flow with squared loss: , . Very wide: .

In the NTK regime the weights barely move, the tangent features stay fixed and the network behaves like a random feature model or kernel method. NTK theory explains optimization of very wide networks, not feature learning in practical ones.

Appears in