TL;DR

  1. Why explain: black-box models are hard to debug and to trust; explanations should help with debugging, trust, the balance of power between provider and customer, recourse, scientific insight, and regulation. The gold standard is an interpretable model (decision tree, linear model, GAM ).
  2. Local post-hoc feature attributions explain one decision afterwards by an extra algorithm: which features mattered? Dimensions: global vs. local, explain vs. the data, model-specific vs. agnostic, data vs. feature attribution, cooperative vs. adversarial. Correlated features and interactions make “importance” ill-defined.
  3. Counterfactual explanation: the smallest change of the input that flips the decision (“with 500 € more income you would get the credit”). Not necessarily actionable or robust.
  4. SHAP: . Observational value function vs. interventional . For a GAM, interventional SHAP recovers the component functions (Theorem 1: SHAP for GAMs). With interactions and dependencies SHAP can be very misleading: “avoid SHAP whenever possible”.
  5. LIME: sample around , fit a local linear surrogate on binned (tabular) or superpixel (image) features. Explanations can be manipulated (off-manifold changes, fooling LIME and SHAP), different algorithms disagree, informative explanations only exist for simple functions. Proposal: explanation cards; for high stakes use interpretable models.
  6. Law: the EU AI Act regulates applications by risk (prohibited, high, limited, minimal; GPAI with systemic risk above FLOPs). GDPR: lawful basis and consent, data minimization, purpose limitation, right to be forgotten (machine unlearning). Copyright: training data needs a legal basis; no copyright without human authorship.

Exam relevance

  • Neither the real exam nor the mock asked about explainability or law, but the lecturer may. Likely forms: multiple choice (what SHAP computes, observational vs. interventional, what LIME fits, AI Act risk categories, GDPR rights) and small computations.
  • Computations: SHAP values by hand for two binary features, a counterfactual for a linear score. Practice with the SHAP widget and the AI Act quiz.
  • Sheet 11, Exercise 1 (counterfactual explanations and algorithmic recourse), Exercise 3 (SHAP explanations and what they do not mean).
  • Everything in one go: the cheat sheet with two full integration tasks at the end of this note.

Overview: 1. Why and what kind of explanations, 2. Issues: correlations and interactions, 3. Counterfactual explanations, 4. SHAP, 5. LIME, 6. Criticism and outlook, 7. Legal issues (AI Act, GDPR, copyright, outside the EU).

(Literature: Christoph Molnar, Interpretable Machine Learning, online, as an introductory textbook.) Slide 75

Explainability for tabular data

Why explainable machine learning?

Slides 77-78

  • Black-box models with millions of parameters are hard to debug and hard to trust.
  • We might want to “understand” what the model does: in science, in critical applications (medicine), in society (loans, university administration).
  • AI regulation requires transparency.

Explanations of why the model predicts something might serve to: debug code; establish trust (society, medicine); balance power and control between model provider and customer; give hints for concrete actions (e.g. recourse); reveal scientific insights (truth?); comply with regulation (EU AI Act, see below).

Setup and the gold standard

Slides 79-80

Setting of this section: ML on feature-based (tabular) data, not text or images, e.g. patients described by numeric values, loan applicants. The data is high-dimensional, the task classification or regression.

Interpretable ML (the gold standard)

Use models so simple that we can understand and judge their decision mechanism:

  • decision trees,
  • linear models ,
  • generalized additive models (GAMs) .

But simple interpretable models might not be powerful enough for complex prediction problems.

Hand-drawn decision tree: blood pressure low or high, then smoker or injured, leading to four treatments
Slide 80: a decision tree is interpretable as a whole.

Local post-hoc explanations and feature attributions

Slides 81-83

Local post-hoc explanation

The model is too complex to understand as a whole. Instead we explain one particular decision with a separate explanation algorithm: it receives a point and the prediction , has white-box access to or can at least sample and query , and tries to justify the decision “post hoc”. The explanation often is a feature attribution: which input features were important for this decision?

One way to build such explanations: if small changes in feature of change the decision about , feature is important. Mathematically this is the gradient , selecting the features with the largest absolute values: gradient-based explanations.

Sketch of a complex decision boundary with a highlighted point, an arrow asking why f of x, and a bar chart of importances for features f1 to f5
Slide 82: a feature attribution for one decision, as a bar per feature.
Two classes separated by a winding boundary; a yellow point near the boundary where moving along feature 2 changes the class but moving along feature 1 does not
Slide 83: moving the yellow point along feature 1 keeps its class, along feature 2 it changes: feature 2 is more important.

Kinds of explanations

Slides 84-88

vs.
global: understand the decision mechanism as a wholelocal: explain only decisions for individual points
explain the decision function (e.g. credit decisions), the data does not matterexplain the data-generating process: relate data and decisions, e.g. find a causal reason for a disease (science)
model-specific: depends on the architecture, e.g. of a neural networkmodel-agnostic: SHAP, LIME, Anchors, …; no strong assumptions on
data attribution: which training points were important for ?feature attribution: which features of were important for ?
cooperative: provider and recipient share a goal (debugging, scientific discovery)adversarial: different goals, e.g. a loan application; the bank has no incentive to give “true” explanations if the customer may use them to sue

Engineering vs. science of explainable AI

Slides 89-92

One can engineer many heuristics that produce explanations, and many of them “make sense”. But what scientific guarantees can we give? When is an explanation “true”? By which standard should we evaluate explanations? Which guarantees hold under which assumptions?

  • The scientific approach is important for the AI Act (e.g. credit scoring), medicine and other critical applications.
  • An engineering approach is justified for exploratory data analysis (science), code debugging, and whenever a useful explanation helps but even many wrong explanations do not really hurt.
Sketch: the true world is mapped with incomplete knowledge to the algorithm's world of points and a decision boundary; from there, ambiguity leads to several different possible feature attribution bar charts; can we reach the true world? rather not
Slide 92: can we assess that an explanation is true? The algorithm only sees an incomplete picture of the world, and many different explanations fit it: rather not.

Local post-hoc explanations (feature attributions)

General issues: correlations and interactions

Slides 94-98

Before the methods, two issues that come up in many of them:

Correlated features in the data domain

Methods that isolate individual features as “reasons” get distorted if features are correlated in the input domain. Examples: the same feature twice (extreme case), years of full-time employment and gender, address and income. The model might pick one of them and ignore the others, use both with different weights, or even use both with opposite weights that cancel in the decision function but not in the gradient.

Interactions in the decision function

Features can also be coupled in the output: they interact. Examples: the XOR function (which of the two features is more important? impossible to answer, we need both), two genes that only jointly trigger a disease. More generally there are -th order interactions of features. What a single importance score should do here is very unclear and may depend on tiny implementation details.

(More: Barocas, Selbst, Raghavan, “The hidden assumptions behind counterfactual explanations and principal reasons”, FAccT 2020.)

Counterfactual explanations

Slides 99-103

(Original paper: Wachter, Mittelstadt and Russell, “Counterfactual explanations without opening the black box: automated decisions and the GDPR”, 2017.)

Counterfactual explanation

What is the smallest change of the input such that the decision function returns a different result? Examples: “If your income were 500 Euro higher, you would get the credit.” “If your gender were male, you would be admitted to the program.” (?!?) One tries to find one or few features whose increase, decrease or change moves the point to the other side of the decision surface.

Sketch in the plane of feature 1 and feature 2: a winding decision surface, a red point to be explained, and a green arrow along feature 1 to the other side
Slide 101: the counterfactual explanation says: change feature 1 to move most quickly to the other class.
Three panels over hours worked and education: a counterfactual arrow from 20 to 41 hours; the box-shaped decision regions a recipient imagines; the fragmented regions of a random forest where the counterfactual is not stable
Slide 102: (a) the counterfactual the recipient gets, (b) the simple, monotone boundary they imagine, (c) how a random forest's boundary may actually look. Counterfactuals of complex functions might not be robust.

Are counterfactuals helpful for recourse?

Counterfactuals are often proposed for recourse: which features to change to get accepted. But:

  • “You would get the credit if you were 10 years younger.” Not actionable.
  • “You would get the credit if you earned 10,000 Euros more per year.” By then you might be 5 years older, and it is unclear whether the counterfactual still holds when the age changes too. And the bank might have retrained its model in the meantime.
  • For complex decision functions the recipient’s intuitive reading (monotone, stable around the counterfactual and the input) may simply be wrong (slide 102).

A counterfactual may give a hint for recourse, but not the full picture.

Task: counterfactual explanation for a linear score

Exam-style task: counterfactuals (4 P)

A bank accepts a credit application iff (income and debts in thousand Euro, age in years). Applicant: income 40, debts 3, age 30.

(a) (1 P, easy) Compute the score and the decision.

(b) (1.5 P, harder) Give the single-feature counterfactuals (change only one feature). Which one is “smallest”? Compare raw units with changes measured in standard deviations (income 15, debts 2, age 12).

(c) (1.5 P, transfer) Discuss the counterfactuals as recourse. What changes if the bank used a random forest instead of the linear score?

SHAP

Intuition from cooperative game theory

Slides 104-106

(Original paper: Lundberg and Lee, “A unified approach to interpreting model predictions”, NeurIPS 2017; thousands of follow-up papers.)

Cooperative game theory evaluates how much influence each player has on the payoff of a coalition (Shapley values). In ML: how much does each feature contribute to the model prediction?

Vanilla idea: the payoff is the prediction at a given point, the players are the features, is the set of all features. To evaluate feature , consider all subsets of features and compare the “prediction based only on the features in ” with the “prediction based on ”: . Sum this over all subsets (with weights) to get the importance of feature .

Definition of SHAP

Slides 107-108

SHAP value

The SHAP value of feature of at is

with a value function (below). The weights sum to 1 over all . The values add up to (efficiency).

Weights for : . For : for , for each of the two , for .

This gives importance scores for all features, plotted as a bar chart, and we “select the important ones” (at least, this is the story).

Bar chart of SHAP feature attributions for features F1 to F12; F1 about minus 0.4, F6 about plus 0.27, the others small
Slide 108: SHAP attributions of one decision as a bar chart.

Memory aid

SHAP = the average extra contribution of feature over all orders in which the features could join the team.

Illustration with pizza

Slides 109-111

  • Which toppings are most influential for the taste of a particular pizza? There are possible toppings; a pizza is a binary vector saying which toppings are on it. Its taste is a number between 1 and 10, say.
  • For a set of toppings, are the values of at the coordinates (e.g. , gives ).
  • The value function rates the contribution of the toppings to the taste: look at all pizzas that agree with on , , and average their taste: . (Tricky detail: which average exactly? See observational vs. interventional.)
  • Importance of topping : compare for all and sum as in the SHAP formula.
A pizza written as a vector p equals p1 to pd, with arrows labeling the coordinates cheese, tomato, salami, pepperoni, mushrooms, and a drawing of a pizza
Slide 109: a pizza as a binary vector of toppings.

Observational vs. interventional SHAP

Slides 112-116

Observational value function

Fix and . Average over all points whose -features coincide with those of :

( the complement of ). Easy in theory, super difficult in practice: estimating the conditional distributions needs a huge amount of data. Pizza: for and , average the taste over all pizzas on the menu with mushrooms and ham. If no pizza on the menu has mushrooms, ham and pineapple, such a pizza does not enter the average.

Interventional (marginal) value function

Set to and sample the remaining features from their marginal distribution: all dependencies between and are broken. “Interventional” because we intervene with the do-operator. Pizza: take all existing pizzas and force mushrooms and ham on them; a pizza with pineapple enters the average with mushrooms and ham added, even if such a pizza is not on the menu (out of distribution).

Which one?

Pretty much everybody uses interventional SHAP: sampling is simple. But it evaluates at out-of-distribution points and breaks dependencies, which may be questionable. If the features are independent, both value functions coincide.

Memory aid

Observational explains the data (conditions on what is known), interventional explains the model (forces values, even unrealistic combinations).

SHAP recovers GAM components

Slides 117-121

A GAM is with component or shape functions , typically non-linear (small trees, low-order polynomials).

Theorem 1 (SHAP for GAMs)

Let be a GAM with centered components, . Then the interventional SHAP values are the component functions:

Without centering, . For observational SHAP the same holds if the features are independent (then both coincide).

What it means in practice: SHAP values are meaningful if the function is a GAM and you use interventional SHAP. They might be misleading if the function is not a GAM. Similar relations hold for higher-order GAMs and higher-order SHAP values (Bordt and von Luxburg, “From Shapley values to generalized additive models and back”, AISTATS 2023).

Task: SHAP values by hand

Exam-style task: SHAP for two binary features (4 P)

Math: (conditional) expectation · SHAP weights

independent with , model , explain with interventional SHAP.

(a) (1 P, easy) Compute , , and .

(b) (1.5 P, harder) Compute and and check efficiency. How is the interaction term split?

(c) (1.5 P, transfer) New model , and in the data is an exact copy of (). Compute interventional and observational SHAP at and interpret.

Practice: SHAP values by hand

Choose a model (coefficients of all terms) and a data distribution over the eight binary points, pick x, and switch between interventional and observational SHAP. The table lists every subset S with its weight and value difference, so each step of the formula can be checked.

Presets: the task example, a GAM (Theorem 1: SHAP for GAMs), XOR, an unused duplicate feature and a duplicate with opposite weights (slide 96).

SHAP with interactions and dependencies

Slides 122-127

  • Interactions (California housing): longitude and latitude alone are not very informative, but jointly they tell whether a house is at the beach, which makes it much more expensive. A GAM cannot describe this; it needs an interaction term. Yet SHAP attributes a medium importance to both and ignores the interaction.
  • Dependencies: a separate feature “ocean proximity” can be replaced by longitude and latitude; the variables are highly dependent. SHAP tends to distribute the importance over all dependent variables, which makes each individual value pretty meaningless.
  • Rankings can be meaningless: because of both effects, truly important features can get a low rank because they are correlated with others. Picking the features with the highest SHAP scores is not a good idea for feature selection.
  • Instead one can estimate the effects separately: standalone contribution, dependencies, interactions (DIP decomposition; König, Günther, von Luxburg, “Disentangling interactions and dependencies in feature attribution”, AISTATS 2025). Wine data: the original scores suggest density and residual sugar are most relevant and citric acid irrelevant. The decomposition shows that residual sugar matters through cooperation: sugar and density are positively correlated (sugar increases density) but have opposing effects on quality that cancel unless both are observed.
Bar chart for the wine quality data: for each feature a standalone part in gray, dependencies and interactions in purple and green, and the original score as a black line
Slide 125: DIP decomposition of feature scores on the wine data: standalone contribution (gray), dependencies (purple), interactions (green).

The lecturer's conclusion: avoid SHAP whenever possible

  • SHAP is very widely used because it has a nice story (game theory) and easy code with nice plots.
  • SHAP is nice for GAMs, but for GAMs we might not need local interpretability methods.
  • In general SHAP can be super misleading, and this is not easy to fix; the literature is full of misinterpretations.
  • Unless you really know what you are doing, avoid SHAP. Many criticisms (correlations, interactions) also concern other feature attribution methods.

LIME

LIME on tabular data and images

Slides 128-133

(Ribeiro, Singh, Guestrin, “Why should I trust you? Explaining the predictions of any classifier”, SIGKDD 2016; Garreau and von Luxburg, “Explaining the explainer: a first theoretical analysis of LIME”, AISTATS 2020; Bordt, Upadhyay, Akata, von Luxburg, “The manifold hypothesis for gradient-based explanations”, 2022.)

LIME (Local Interpretable Model-agnostic Explanations)

We cannot replace a complex function (random forest, deep network) globally by a simple function on few “explainable” features (it would not be accurate). But we might approximate it locally by a simple linear function to understand the decision at , hoping that this local approximation is meaningful. Rough sketch for tabular data ():

  1. Fix the point whose decision should be explained.
  2. Sample points locally around , evaluate , apply a binning procedure to the features.
  3. Fit a simple linear function that approximates locally.
  4. Use the few most prominent coordinates of this linear model as the explanation.

LIME only needs to evaluate at arbitrary inputs: it works for any black box.

A curved function f in blue with a red tangent line at the point xi; below, sampled points with weights pi shown as vertical bars and bin boundaries q in green
Slide 131: LIME on tabular data. Samples around the point (weights as vertical bars) are binned (green boundaries), and a linear model (red) approximates f locally.

Images: pixels are not interpretable features. The image is split into superpixels (contiguous patches, e.g. ) in a pre-processing step. LIME samples images near the given one by randomly switching superpixels on and off; the binary on/off vector plays the role of the feature vector. It then identifies the superpixels that contribute most to the classification.

An image of a strawberry and a tortoise, its superpixel segmentation, and the superpixels that explain the classes terrapin and strawberry
Slide 132: LIME on an image. Superpixels explaining "terrapin" and "strawberry".

Criticism and outlook

Explanations can be manipulated

Slides 134-137

(Anders et al., “Fairwashing explanations with off-manifold detergent”, ICML 2020; Slack et al., “Fooling LIME and SHAP: adversarial attacks on post hoc explanation methods”, AIES 2020; Slack et al., “Counterfactual explanations can be manipulated”, 2021. More general: Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead”, Nature Machine Intelligence 2019; Bordt, Finck, Raidl, von Luxburg, “Post-hoc explanations fail to achieve their purpose in adversarial contexts”, FAccT 2022.)

How manipulation works

Given , construct that makes the same decisions on the data but behaves very differently off the data distribution. Real data often lives on a low-dimensional manifold; around each data point there are many off-manifold points. Explanation algorithms sample from a neighborhood of , often off-manifold (or use gradients). By changing off-manifold we change the explanation at without changing any decision.

Grid of FashionMNIST inputs with gradient, input times gradient, integrated gradients and LRP explanations; the manipulated model always shows the digits 42
Slide 136: manipulated gradient-based explanations always show "42" (Anders et al.).
Two SHAP force plots: the biased classifier's explanation is dominated by race; after the attack, two uncorrelated features appear most important
Slide 137: a classifier that heavily uses "race" gets an innocent-looking SHAP explanation after the attack (Slack et al.).

Different algorithms, different explanations

Slides 138-141

Four bar charts of feature attributions for the same decision: SHAP highlights F1 and F6, LIME F9, DiCE F9 positively, interventional SHAP F1, F4, F6 and F9
Slide 138: SHAP, LIME, DiCE and interventional SHAP explain the same decision of the same model completely differently (Bordt et al., FAccT 2022).
  • In an adversarial setting the “opponent” can cherry-pick the explanation they like; for an examiner this is pretty much impossible to detect. We cannot trust these explanations if we don’t trust the one who invokes the algorithm.
  • Informative explanations only exist for simple functions: in a framework for when a local explanation is “informative”, one can prove this, questioning the whole point of local post-hoc explanations (Günther, Szabados, Bhattacharjee, Bordt, von Luxburg, arXiv 2025).
  • Explanation cards (Günther, Szabados, König, Meding, Bordt, von Luxburg, arXiv 2026), the constructive way forward: augment explanations with information about robustness and validity and clear instructions for interpretation. This can make otherwise uninformative explanations useful, helps detect when they are not, and shifts responsibility from users to providers, who must state upfront what can and cannot be concluded.

The lecturer's summary

If explanations are designed for a different end user, supply explanation cards. If reliable explanations matter (e.g. in a societal context), use interpretable algorithms, do not allow complicated models with post-hoc explanations. Other people have different opinions here.

Outlook: images, language models, mechanistic interpretability

Slides 142-144

  • Image classifiers: many methods, many based on the gradient of the output with respect to the input pixels (which pixels affect the prediction most): saliency maps (Simonyan et al., ICLR Workshop 2014), integrated gradients (Sundararajan et al., ICML 2017), Grad-CAM (Selvaraju et al., ICCV 2017), SmoothGrad (Smilkov et al., 2017).
  • Language models: people ask LLMs to explain themselves or their reasoning. The lecturer is not convinced, and it is definitely not trustworthy when the explanation really matters. A vastly open field.
  • Mechanistic interpretability: try to extract what sub-components of a trained network do, e.g. which sub-task a particular attention head performs. Some small successes, but big questions: does it generalize beyond toy applications, and how trustworthy are the findings?

Legal issues

Why regulate, and how legislation works

Slides 147-148

Reasons for regulating AI systems:

  • Safety and harm prevention: AI can cause accidents, discrimination etc. at scale.
  • Accountability: who is liable when an AI system causes harm?
  • Fundamental rights: privacy, non-discrimination.
  • Market integrity: consumer protection, a level playing field.
  • Trust: public acceptance of AI requires oversight and legal certainty.

How legislation works: law sets norms and principles at a high level, not technical specifications. It does not define how to comply; that is left to standards bodies, industry guidelines and case law. Many key concepts are deliberately vague (what does “transparent” mean for an ML model?), and their interpretation is refined over time by court rulings and national regulators. Compliance is an ongoing process, not a one-time checklist.

The EU AI Act

Slides 150-155

  • The world’s first comprehensive AI regulation. It applies to providers, deployers and importers of AI systems placed on the EU market or affecting persons in the EU.
  • Goals: AI that is safe, transparent and non-discriminatory; innovation through legal certainty.
  • Timeline: proposed April 2021, political agreement December 2023, in force August 2024, stepwise roll-out until August 2026.
  • Reading it: the full text starts with the non-binding recitals (background, motivation, aims, interpretation); the legally binding part starts in the middle (“Chapter I: General provisions”).

Risk-based approach

The AI Act does not regulate “algorithms” but their applications, sorted into risk categories:

  • Unacceptable risk (prohibited): real-time biometric surveillance in public spaces, social scoring by governments, subliminal manipulation, exploitation of vulnerable groups.
  • High risk: safety-critical systems (medical devices, critical infrastructure, autonomous vehicles) and high-impact systems (credit scoring, recruitment, education, law enforcement, border control).
  • Limited risk (chatbots, deepfakes) and minimal risk (recommender systems, spam filters, games): little or no mandatory obligations.

The major part of the AI Act is about high-risk systems.

Obligations for high-risk systems

Risk management over the entire lifecycle; data governance (high-quality training, validation and test data, documented data practices); technical documentation and automatic logging; transparency toward deployers and instructions for use; human oversight (humans can intervene or override).

General-purpose AI (GPAI) models

Added only in the last round of debates, because ChatGPT entered the market. Article 3(63): “an AI model that is trained with a large amount of data using self-supervision at scale, that displays significant generality and is capable of competently performing a wide range of distinct tasks regardless of the way the model is placed on the market and that can be integrated into a variety of downstream systems or applications, except AI models that are used for research, development or prototyping activities before they are placed on the market.”

  • Systemic risk threshold: training compute FLOPs (roughly GPT-4 scale and above).
  • All GPAI providers: technical documentation, comply with copyright law, publish a summary of the training data.
  • Systemic-risk models additionally: adversarial testing / red-teaming before release, reporting serious incidents to the European Commission, cybersecurity measures, energy efficiency reporting.

Practice: which risk category?

Fourteen applications written for these notes. Pick the risk category as the lecture defines it; the explanation points to the slide. The legal text has exceptions and details beyond the lecture.

GDPR

Slides 157-159

  • The General Data Protection Regulation: EU regulation in force since May 2018, regulating the processing of personal data of natural persons. It applies to any entity processing personal data of EU residents, wherever the entity is located. Relevance for ML: training data frequently contains personal data.
  • Data consent: lawful reasons for processing are consent, contract, legal obligation, legitimate interest, … Special categories (health, biometrics, ethnicity, political opinions, …) need explicit consent. Data minimization: collect only what is necessary for the stated purpose. Purpose limitation: data may not be reused beyond the original purpose. ML implication: scraping publicly available data does not automatically grant permission to use it for training.
  • Right to be forgotten: individuals can request the erasure of their personal data. Once data is encoded in model weights, deletion is non-trivial; active research area machine unlearning: removing the influence of specific training points from a model.

Slides 161-162

  • Training data (EU Copyright Directive 2019): using any data needs a legal basis: a license, public domain or open licenses (Creative Commons), own data, a data sharing agreement, … GDPR applies additionally for personal data. Scraping without permission is increasingly contested, e.g. New York Times v. OpenAI, Getty Images v. Stability AI.
  • Ownership of AI output: current consensus: no copyright without human authorship; AI has no legal personhood.
  • Training as copyright infringement: an active legal dispute, not yet settled by courts in the US or the EU.
  • Trade secrets: model weights and training pipelines can be protected as proprietary IP.
  • Patents: an AI cannot be listed as inventor.

Regulation outside the EU and take-aways

Slides 164-165

  • UK: no dedicated AI law; existing sector regulators cover many cases.
  • Canada: relies on existing privacy, consumer protection, human rights and sector rules (an attempt at AI regulation was stopped in 2025).
  • USA: no federal AI law; sector-specific rules (medical AI, consumer protection, discrimination, …).
  • China: Algorithmic Recommendation Provisions (2022), Generative AI Regulation (2023); emphasis on content control and alignment with state interests.
  • No global consensus on regulation.

Key take-aways

  • Know your data: legal basis, provenance, personal data.
  • Know your system: its risk category under the AI Act and the resulting obligations.
  • Know your outputs: generated content carries copyright risk for the deployer.

The legal landscape is evolving rapidly; landmark court decisions and implementing regulations are still pending.

Summary

TopicKey message
why explaindebugging, trust, power balance, recourse, science, regulation; gold standard: interpretable models (trees, linear, GAMs)
local post-hocan extra algorithm explains one decision , usually as feature attributions (e.g. gradients)
kindsglobal/local, function/data, model-specific/agnostic, data/feature attribution, cooperative/adversarial
issuescorrelated features and interactions make importance ill-defined
counterfactualssmallest change that flips the decision; depends on the metric, not always actionable or robust
SHAPShapley-weighted average of ; observational vs. interventional ; recovers GAM components; misleading with interactions and dependencies
LIMElocal linear surrogate from samples around (bins, superpixels)
criticismmanipulation off-manifold, algorithms disagree, cherry-picking, informative only for simple functions; explanation cards
AI Actapplications by risk: prohibited, high (obligations), limited, minimal; GPAI, systemic risk FLOPs
GDPR, copyrightlawful basis and consent, minimization, purpose limitation, right to be forgotten (unlearning); training data needs a legal basis, no copyright without a human author

Self-Test

Multiple Choice

Cheat sheet and full integration tasks

The first task covers the two calculations of this lecture, SHAP values and a counterfactual for a linear score. The second is one scenario in which the terms on explanation methods, the AI Act and the GDPR have to be applied once. Write your own sheet first, solve the tasks with it next to you, then open the sheet at the bottom and compare. The letters in brackets name the block of the sheet that a subtask needs.

Full integration task: SHAP and a counterfactual by hand (14 P)

Math: (conditional) expectation · SHAP weights

Part 1. independent with . The model is . Explain the point with interventional SHAP.

(a) (3 P, block A) Compute , , and .

(b) (3 P, block A) Compute and and check efficiency. What does the sign of say?

(c) (2 P, block A) Drop the interaction term, . Give the SHAP values of at without a calculation over subsets. Which part of and in (b) comes from the interaction?

Part 2. A bank accepts a credit application iff (income and debts in thousand Euro). Applicant: income 35, debts 4, 10 years employed.

(d) (3 P, block B) Compute the score and the decision. Give the counterfactual for each single feature, if one exists.

(e) (3 P, block B) Which single-feature counterfactual is smallest when changes are measured in standard deviations (income 10, years employed 8)? Give a counterfactual that changes two features and uses the debts. Discuss both as recourse.

Full integration task: one credit model, every term (10 P)

A bank uses a random forest for credit decisions. Rejected applicants receive a bar chart of SHAP values as the explanation.

(a) (2 P, block C) Classify this explanation: global or local, model-specific or model-agnostic, data or feature attribution, cooperative or adversarial setting?

(b) (2 P, blocks A and C) The features “address” and “income” are strongly correlated. What does that do to the SHAP values, and where do observational and interventional SHAP differ?

(c) (2 P, block C) How could the bank change the explanations without changing a single decision? What does the lecture recommend where reliable explanations matter?

(d) (2 P, block D) Which risk category of the EU AI Act does credit scoring fall into, and which obligations follow? Give the category of a customer service chatbot, of social scoring by a government and of a spam filter.

(e) (2 P, block E) An applicant asks the bank to delete their personal data, which was used to train the model. And the bank had collected public social media profiles as extra training data. What does the GDPR say to both?

References