Definition
Among the many minimizers of an over-parameterized objective, the optimizer (GD, SGD) converges to one with extra properties, without an explicit penalty term.
Formula
Least squares, , full row rank, GD from with :
Logistic loss, separable data: and (max margin).
The iterates stay in the span of the data, so the component in is never touched: from GD converges to the solution closest to . For deep non-linear networks the implicit bias is not well understood.
Appears in
- Lecture 2, SGD picks a special minimizer
- Lecture 8, Theorem 1 (GD finds the minimum norm solution), Theorem 2 (max-margin bias of GD), neural networks