A knowledge base is an organized repository of knowledge about a domain, maintained so that people or software can retrieve and use it. The term has two principal meanings: a collection of explanatory articles used in organizational support, and a formally represented collection of facts and rules used in artificial intelligence. These meanings overlap in their emphasis on reusable information, but differ in how content is represented, queried, and interpreted. A knowledge base may therefore be a searchable help center, a logical theory, or an information source supporting an automated application. (confluence.atlassian.com)
Scope and related concepts
In organizational settings, knowledge bases contain frequently asked questions, procedures, troubleshooting instructions, and explanations of known problems. They can serve external customers or provide a restricted workspace for employees. Compared with general software documentation, a support knowledge base often emphasizes particular incidents, unusual cases, and solutions developed through operational experience. Its articles may be consulted independently or supplied by support staff during an interaction. (atlassian.com)
In knowledge representation and reasoning, the repository instead contains statements expressed in a language with defined semantics. A system can use those statements and additional observations to determine what follows from them. Here, “knowledge” means information asserted within the representation, rather than a guarantee that every assertion accurately describes reality. Correctly interpreting an answer requires understanding what the repository’s symbols denote. (cs.ubc.ca)
A database can provide the storage infrastructure for a knowledge base, but the concepts are not identical. “Database” emphasizes organized data storage and access; “knowledge base” emphasizes the information’s role in consultation or reasoning. An ontology-oriented repository may use database technology while applying assumptions about missing information that differ from those common in database applications. (w3.org)
Representation and organization
Human-oriented repositories commonly organize articles through titles, topic labels, templates, and navigation pages. Templates give procedures or troubleshooting explanations a consistent structure, while labels associate content with relevant topics. Information retrieval and topic navigation let readers locate material without necessarily traversing a fixed hierarchy. Notifications, comments, and article feedback provide mechanisms for communicating changes and identifying unclear content. (confluence.atlassian.com)
Machine-oriented repositories represent entities, properties, relationships, and general constraints. An ontology establishes vocabulary and describes relationships among its terms. It may include both general statements about classes and assertions concerning individual objects. Ontology engineering consequently involves decisions about what distinctions matter within the domain and how those distinctions are expressed formally. (w3.org)
The Resource Description Framework (RDF) represents statements as subject–predicate–object triples. A collection of triples forms a graph, connecting resources through named relationships. Resources can have globally identifiable names, facilitating the combination of information from different sources. RDF provides a representation model, while languages such as the Web Ontology Language (OWL) add constructs for expressing richer relationships and supporting inference. Graph-shaped representation does not, by itself, establish the truth or completeness of the represented information. (w3.org)
Reasoning and logical interpretation
Within symbolic artificial intelligence, a knowledge base supplies explicitly represented information to a reasoning procedure. Languages based on propositional logic or first-order logic can express facts and general rules. Through deductive reasoning, a system answers questions by deriving consequences rather than merely locating a stored sentence. (cs.ubc.ca)
For example, suppose a repository states that every refrigerated shipment requires temperature monitoring and that shipment A is refrigerated. It follows that shipment A requires temperature monitoring. This illustrates logical consequence: the conclusion holds in every interpretation in which the premises hold. It does not establish independently that the premises are accurate descriptions of the shipment. (artint.info)
Treatment of missing information is especially important. Under the open-world assumption, an unstated proposition is not automatically false; its truth may be unknown. OWL follows this approach. Under the closed-world assumption, information not established within the relevant repository is treated as false. These assumptions produce different answers to questions about incomplete records and must not be confused with simple differences in storage format. (w3.org)
Knowledge bases in language-model systems
A knowledge base can supply external information to a large language model through retrieval-augmented generation (RAG). A retrieval component selects relevant passages, which the generation component uses when producing an answer. The original RAG research combined a pretrained sequence-to-sequence model with a dense index of Wikipedia passages, distinguishing information stored in model parameters from information accessed through an external repository. (arxiv.org)
Document-based implementations commonly divide material into passages, index it, and retrieve candidates for a query. Their pipelines can include query transformations and reranking before generation. This makes the repository distinct from training data used to learn model parameters. Retrieval allows updated or domain-specific information to be supplied without requiring every change to be incorporated into those parameters. Nevertheless, retrieval quality and generation quality remain separate issues: relevant documents can be missed, and retrieved evidence does not guarantee a correct answer. (arxiv.org)
Publication, access, and evaluation
Maintaining a support repository involves more than storing completed articles. Drafting and publication controls distinguish work in progress from material presented to readers. Permissions separate internal information from customer-facing content, while connections to service workflows allow solutions discovered during support work to become reusable articles. Restricted pages can remain available to staff without appearing in a public help center. (support.atlassian.com)
Evaluation depends on purpose. A support knowledge base can be assessed through article views, helpfulness feedback, sharing, and requests resolved or avoided through articles. A formal repository can be examined for consistency and the consequences of its axioms; a retrieval-based application also requires assessment of retrieval and generated responses. These measurements address different properties, so neither a large article count nor successful logical inference alone demonstrates that a repository meets its users’ information needs. (atlassian.com)