Explainable artificial intelligence (XAI) is a field of artificial intelligence concerned with making the behavior, outputs, and limitations of computational systems understandable to people. It includes both models whose operation can be inspected directly and techniques that explain otherwise opaque systems. In machine learning, explanations may identify influential inputs, describe decision rules, or show how different inputs would change a prediction. An explanation must be distinguished from a persuasive justification: being understandable does not necessarily mean accurately representing the system’s operation. (nvlpubs.nist.gov)
Scope and distinctions
Interpretability and explainability overlap, but their definitions vary across research communities. Interpretability often concerns whether a person can understand a model’s structure or behavior; explainability often concerns the provision of reasons or evidence for an output. Both depend on the audience and task. A developer investigating an error, an auditor reviewing a system, and a person affected by a decision may require different information. (nvlpubs.nist.gov)
A central distinction separates intrinsically interpretable models from post-hoc explanations. Small decision trees, sparse rule lists, and suitably constrained linear regression models can expose their prediction mechanisms directly. Post-hoc methods instead analyze a trained model without necessarily making its complete operation understandable. Complexity matters: a very large tree or a linear model containing many transformed variables may remain difficult to interpret. There is no universal rule that greater predictive accuracy requires lower interpretability. (nature.com)
Explanations can also be local, addressing one prediction, or global, describing behavior across many inputs. A local explanation does not establish that the same relationship holds everywhere. Model-agnostic methods operate through inputs and outputs, whereas model-specific techniques use properties such as internal structure or accessible derivatives. (arxiv.org)
Principal methods
Local surrogate models. LIME, introduced in 2016, approximates a model near a selected input using a simpler explanatory model. It generates perturbed examples, obtains predictions, and fits an interpretable approximation weighted toward the input being explained. The resulting explanation is local; its usefulness depends on the sampled neighborhood, interpretable representation, and agreement between surrogate and original model. (arxiv.org)
Feature attribution. Attribution methods distribute a prediction or prediction difference among input features. SHAP, introduced in 2017, provides a framework for additive explanations drawing on Shapley values from game theory. Features are treated as participants contributing to a prediction relative to a specified reference. Attribution depends on how unavailable features and background data are represented, so values are not context-free measures of importance. (proceedings.neurips.cc)
Gradient-based explanations. For differentiable neural networks, a gradient measures local output sensitivity to an input. Integrated gradients accumulates gradients along a path from a baseline input to the actual input. This method was developed around formal attribution requirements, including sensitivity and implementation invariance. Its interpretation remains tied to the selected baseline and output quantity. Applications include computer vision and natural language processing. (proceedings.mlr.press)
Counterfactual explanations. A counterfactual identifies changes to an input that would produce a different model outcome. For example, an illustrative explanation might specify what changes would move a prediction across a classification threshold. Such explanations need not reveal the full model. They describe an alternative under the model, however, rather than guaranteeing that a corresponding real-world action is feasible or would produce the same result. (arxiv.org)
Evaluation and reliability
The US National Institute of Standards and Technology’s 2021 report identifies four principles: explanation, meaningfulness, explanation accuracy, and knowledge limits. Respectively, these concern providing reasons or evidence, making explanations understandable to their intended recipients, correctly reflecting the process generating an output, and recognizing the conditions within which a system can operate appropriately. Explanation accuracy is distinct from prediction accuracy: an incorrect prediction can be explained faithfully, while a correct prediction can receive a misleading explanation. (nvlpubs.nist.gov)
Evaluation therefore requires more than attractive visualizations or user satisfaction. Researchers examine whether explanations actually depend on learned parameters and training data, and whether they track relevant changes in model behavior. In “Sanity Checks for Saliency Maps” (2018), experiments using randomized parameters and labels showed that some visually convincing methods were insensitive to properties they were expected to explain. A plausible-looking heat map is consequently insufficient evidence of faithfulness. (papers.nips.cc)
Human-centered evaluation asks whether explanations help people perform a specified task, such as identifying unreliable predictions or choosing between models. LIME’s original experiments examined these uses rather than treating explanation quality solely as visual appeal. The relevant objective is informed reliance, not simply increasing trust. (arxiv.org)
Uses and limitations
Explanations support model debugging, inspection, and error analysis. They can help reveal reliance on unintended cues and clarify differences between candidate systems. Nevertheless, an approximation may omit important behavior, and an explanation of model dependence does not by itself establish causation in the world. Intrinsically interpretable models and post-hoc explanations offer different forms of access to a decision mechanism and should not be treated as interchangeable. (arxiv.org)
For large language models, fluent verbal explanations present a particular reliability problem. Research on chain-of-thought prompting has demonstrated cases in which biasing prompt features influenced answers but were not acknowledged in the accompanying explanations. Generated reasoning text can therefore be useful material for investigation without constituting a verified account of the computation that produced an answer. Assessing explanation faithfulness requires evidence beyond the model’s own description. (arxiv.org)
References
- Four Principles of Explainable Artificial Intelligencenvlpubs.nist.gov
- Stop explaining black box machine learning models for high stakes decisions and use interpretable models insteadnature.com
- "Why Should I Trust You?": Explaining the Predictions of Any Classifierarxiv.org
- A Unified Approach to Interpreting Model Predictionsproceedings.neurips.cc
- Axiomatic Attribution for Deep Networksproceedings.mlr.press
- Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPRarxiv.org
- Sanity Checks for Saliency Mapspapers.nips.cc
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Promptingarxiv.org