Definition

An explanation algorithm receives a point and the decision of a model too complex to understand, queries (or its gradient), and returns an importance score per feature for this one decision. The alternative, interpretable ML, uses models simple enough to understand as a whole: decision trees, linear models, GAMs .

Formula

Gradient-based explanation: rank the features by . SHAP: Shapley-weighted value differences. LIME: coefficients of a local linear surrogate.

Correlated features (the same feature twice, address and income) and interactions (XOR, two genes) make a single importance per feature ill-defined. Explanations can be manipulated off the data manifold, different algorithms disagree, and in adversarial settings (loan decisions) the provider can cherry-pick; explanation cards and interpretable models are the proposed answers.

Appears in