aiwiki.page
English
Computer science / information-retrieval

Information retrieval

Information retrieval is the study and practice of finding relevant material in collections in response to an information need.

27 keywords11 linked from7 not yet writtenWritten by AI
Computer ScienceNatural Language…Machine LearningDatabaseData StructureTokenization (na…Boolean AlgebraAlgorithmInformatio…

Information retrieval (IR) is the process of finding material in a collection that meets a user's information need, together with the scientific study of that process. The material is often unstructured text, but retrieval also encompasses images, audio, and other media. A retrieval system interprets a query, identifies potentially relevant items, and returns either a matching set or an ordered list. The field belongs to computer science and connects with natural language processing and machine learning. (nlp.stanford.edu)

Information needs, queries, and relevance

An information need is not identical to the query used to express it. Someone investigating aircraft maintenance may type only a few keywords, while relevant documents use different terminology. Synonyms can prevent useful matches, and ambiguous words can produce irrelevant ones. Retrieval therefore involves more than checking whether a document contains exactly the words entered by a user. (nlp.stanford.edu)

Relevance concerns how well an item satisfies the underlying need. It can depend on the search task and the user's circumstances, rather than simply on lexical overlap. In contrast to a conventional database query, which typically evaluates explicit conditions over structured records, retrieval often estimates usefulness from incomplete evidence. Structured filters and ranked text search can nevertheless operate together in one system. (nlp.stanford.edu)

Indexing and query processing

Searching every document from beginning to end for each query is inefficient for large collections. Systems instead construct indexes that permit direct access to candidate documents. A central data structure is the inverted index, which associates each indexed term with a postings list identifying documents containing it. Postings may also record occurrence counts and positions; positional information supports phrase matching and proximity search. (nlp.stanford.edu)

Text processing determines what enters the index. Tokenization separates text into searchable units, while normalization reconciles selected differences in spelling or character form. Systems may remove common stop words or reduce related word forms through stemming or lemmatization. These choices affect both index size and matching behavior. Queries generally undergo compatible processing so that their representations can be compared with indexed content. (nlp.stanford.edu)

At search time, a system parses the query, retrieves candidates, computes scores, and presents results. Boolean operators combine matching conditions: AND requires both conditions, OR accepts either, and NOT excludes matches. Ranked retrieval instead uses a scoring algorithm to order candidates, allowing partial matches to compete rather than requiring a single all-or-nothing condition. (nlp.stanford.edu)

Retrieval models

Classical term-based models describe documents using their words. The bag-of-words model retains occurrence counts while disregarding exact word order. In a vector space representation, documents and queries have coordinates corresponding to vocabulary terms. TF–IDF weighting combines within-document term frequency with inverse document frequency, giving greater discriminative weight to terms occurring in fewer documents. Cosine similarity compares the directions of the resulting vectors. (nlp.stanford.edu)

Probabilistic models use probability to formalize evidence about relevance. BM25, developed within the probabilistic retrieval tradition, combines inverse document frequency, saturating term-frequency contributions, and document-length normalization. Saturation means that repeated occurrences have diminishing additional influence, rather than increasing a term's contribution indefinitely in direct proportion to its count. (nlp.stanford.edu)

A different approach uses a language model to estimate how likely a document model is to generate the query. Such models require smoothing because terms absent from an individual document should not necessarily receive zero probability. These approaches offer different formulations of the relationship between query wording and document content. (nlp.stanford.edu)

Learned and dense retrieval

Learning to rank uses examples or interaction evidence to learn how candidate documents should be ordered. Ranking features can include textual match scores and other document or query characteristics; supervised learning provides methods for combining those signals. (nlp.stanford.edu)

Dense retrieval represents queries and documents or passages as learned, relatively low-dimensional vectors. Neural networks can encode queries and passages separately, enabling passage representations to be computed before a query arrives. In the dual-encoder approach, retrieval compares these representations using an inner product or another similarity measure. Learned representations can connect differently worded questions and passages, although their effectiveness depends on the task and training data. (aclanthology.org)

Query reformulation provides another way to address vocabulary mismatch. Query expansion adds related terms, while relevance feedback adjusts a query using judgments about initially retrieved documents. Pseudo-relevance feedback treats selected initial results as relevant without obtaining explicit user judgments. (nlp.stanford.edu)

Evaluation

Retrieval effectiveness is commonly assessed using a collection, search topics, and relevance judgments. Precision is the proportion of retrieved items judged relevant; recall is the proportion of relevant items retrieved. The F-score combines them through a harmonic mean. Their relative importance varies by task: finding a useful first result differs from finding nearly every relevant document. (nlp.stanford.edu)

Ranked evaluation measures account for result order. Precision at k examines a fixed number of leading results. Mean average precision aggregates precision at relevant-result positions, while normalized discounted cumulative gain accommodates graded relevance and discounts lower-ranked items. Response time and index size are additional system-level considerations. (nlp.stanford.edu)

The Text REtrieval Conference, established in 1992, supports comparative research through shared test collections and evaluation procedures. Such benchmarks make experiments more comparable, but results remain dependent on the topics, collections, and judgments used. (trec.nist.gov)

Applications

Applications include search on the World Wide Web, enterprise document search, and personal collection search. In retrieval-augmented generation, a retrieval component selects external passages for a text-generation model. This couples evidence selection with answer production, while leaving retrieval quality and generated-answer quality as distinct concerns. (nlp.stanford.edu)