AI hallucination is a phenomenon in which an artificial intelligence system generates content that is factually incorrect, unsupported by the available evidence, or inconsistent with its input. It is particularly associated with generative artificial intelligence and large language models, whose fluent responses can make invented information appear credible. The US National Institute of Standards and Technology uses confabulation for overlapping failures, including confidently presented falsehoods and contradictions. The term describes observable output behavior, rather than establishing that a machine has human-like perceptual experiences. (nvlpubs.nist.gov)
Definition and scope
There is no universally accepted operational definition of hallucination. In natural language processing, researchers distinguish factuality, meaning agreement with reliable information about the world, from faithfulness, meaning agreement with the source material or instructions supplied for a task. These criteria can diverge: a summary may faithfully reproduce an inaccurate source, while an independently true statement may be unsupported by the document being summarized. An audit published in 2024 found substantial disagreement about terminology across research publications and practitioners. (aclanthology.org)
A common source-based classification distinguishes:
- Intrinsic hallucination: content that contradicts the source, such as changing a reported quantity or attributing an action to the wrong person.
- Extrinsic hallucination: content that cannot be verified from the source, such as adding an unmentioned event.
Extrinsic content is not necessarily false in the wider world; its defining feature is lack of support within the task’s evidence. Consequently, evaluations must specify whether they assess source consistency, external factual accuracy, or both. (aclanthology.org)
Hallucination research predates conversational assistants and encompasses automatic text summarization, machine translation, dialogue, and visual-language generation. In multimodal systems, the corresponding failure can involve describing objects or events that the visual input does not support. Whether invented material counts as an error depends on the task: imaginative generation and evidence-constrained reporting have different requirements. (arxiv.org)
Mechanisms and contributing factors
A language model learns patterns from training data and generates continuations using a conditional probability distribution. In conventional autoregressive training, a loss function rewards predicting observed tokens, commonly through maximum likelihood estimation. This objective does not independently verify each generated claim. A continuation can therefore be linguistically plausible without being factually justified. Research on abstractive summarization has documented this mismatch between generation objectives and source faithfulness. (nvlpubs.nist.gov)
Incorrect output can arise from learned misinformation as well as invented combinations of otherwise familiar information. The TruthfulQA study demonstrated that models could reproduce common human misconceptions learned through imitation of text. Its results also showed that increasing model size did not automatically improve truthfulness on that particular benchmark; this finding is task-specific, not a universal relationship between size and reliability. (aclanthology.org)
Confidence presents a separate problem. A high probability for a token sequence is not equivalent to a verified probability that its claims are true. Research on model self-evaluation has found useful signals of answer correctness under some conditions, but their generalization and calibration can deteriorate under distribution shift. Experiments also distinguish ordinary next-token predictions from specially elicited judgments about correctness. (arxiv.org)
Detection and measurement
Hallucination is difficult to measure because a single response can mix supported and unsupported claims. Human evaluation can compare output with source documents or authoritative evidence, but judgments depend on annotation instructions, the evidence available, and the unit being assessed. A survey of 64 studies published between 2019 and 2024 identified substantial variation in human evaluation practices. (aclanthology.org)
FActScore, introduced in 2023, addresses mixed factuality by dividing long-form output into atomic facts and calculating the proportion supported by a reliable knowledge source. This yields a more fine-grained measure than labeling an entire passage correct or incorrect. The score measures factual precision relative to the chosen evidence; it does not, by itself, establish completeness or usefulness. (aclanthology.org)
Other approaches assess whether generated claims follow from the input or compare multiple sampled answers. SelfCheckGPT uses agreement and disagreement across samples as a detection signal without requiring an external database. Inconsistent samples can indicate unreliable content, although agreement alone cannot logically establish truth: repeated responses may share the same error. (aclanthology.org)
Benchmark results require careful interpretation. TruthfulQA targets questions likely to elicit misconceptions, whereas FActScore evaluates supported claims in longer responses. Scores from such different tasks are not interchangeable measures of a system’s universal hallucination rate. (aclanthology.org)
Mitigation and limitations
Retrieval-augmented generation combines a generator with information retrieval over external material, such as a knowledge base. The original 2020 RAG study reported more factual generation than its parameter-only comparison system on the evaluated tasks. Retrieved evidence also provides a basis for checking claims and updating accessible information without relying exclusively on knowledge encoded in model parameters. These results demonstrate improvement under tested conditions, not a guarantee that every retrieved or generated statement is correct. (arxiv.org)
Training-based approaches include fine-tuning with objectives that distinguish factual accuracy from imitation. Self-evaluation methods seek to identify answers a model is unlikely to produce correctly, potentially supporting selective answering rather than unrestricted generation. Their effectiveness depends on task design and how well uncertainty estimates transfer to new inputs. (aclanthology.org)
Consequences and system evaluation
Confabulated content can include fabricated citations, misleading explanations, and contradictions that undermine information integrity. Its significance depends on the application and whether downstream processes detect the error. NIST’s generative AI risk profile treats confabulation as a reliability concern and includes testing, documentation, and ongoing monitoring among risk-management activities. Evaluation therefore concerns not only the model’s output, but also the evidence and verification procedures surrounding its use. (nvlpubs.nist.gov)