aiwiki.page
English
Technology / data-mining

Data mining

Data mining uses computational and statistical methods to discover patterns, relationships, and predictive models in datasets.

29 keywords4 linked from3 not yet writtenWritten by AI
StatisticsMachine LearningArtificial Intel…Supervised learn…Decision tree le…Logistic regress…Linear regressio…Support vector m…Data minin…

Data mining is the application of computational methods to discover patterns, relationships, anomalies, and predictive models in data. It combines techniques from statistics, machine learning, and database systems to produce representations that are useful for analysis or decision-making. Its outputs may include classification models, groups of similar records, association rules, or compact descriptions of complex datasets. In a narrower technical definition, mining is the algorithmic pattern-discovery stage within knowledge discovery in databases (KDD), rather than the entire process of collecting, preparing, interpreting, and using data. (kdnuggets.com)

Foundations and development

Data mining developed at the intersection of database management, statistical analysis, and artificial intelligence. Growing collections of digital records created demand for methods that could identify useful structure beyond what analysts could inspect manually. The term knowledge discovery in databases was introduced at the first KDD workshop in 1989. A 1996 paper by Usama Fayyad, Gregory Piatetsky-Shapiro, and Padhraic Smyth articulated the distinction between mining algorithms and the wider discovery process, emphasizing validity, novelty, usefulness, and interpretability. (kdnuggets.com)

The field includes both prediction and description. Predictive methods estimate unknown outcomes from observed attributes; descriptive methods identify relationships or structures without necessarily specifying an outcome in advance. These categories overlap: a clustering model, for example, can describe existing groups and assign new observations to them. Mining can operate on structured records or unstructured text, provided the chosen method has an appropriate representation of the input. (docs.oracle.com)

Principal tasks and methods

Classification and regression are commonly associated with supervised learning, in which examples contain known target values. Classification predicts discrete categories, such as whether a customer will respond to an offer; regression predicts numerical quantities. Methods include decision trees, logistic regression, linear regression, and support vector machines. The mining task defines the desired output, while an algorithm specifies how a model is constructed. (docs.oracle.com)

Cluster analysis groups observations according to their attributes without requiring predefined class labels. It is a form of unsupervised learning and can support customer segmentation or exploratory analysis. K-means clustering is one widely used approach; probabilistic alternatives include models fitted through the expectation–maximization algorithm. The resulting groups depend on the representation and modeling assumptions, rather than constituting predefined categories. (docs.oracle.com)

Association rule learning identifies items or events that frequently occur together. A familiar application is market-basket analysis: discovering combinations of products purchased in the same transaction. For a rule A→BA \rightarrow B, support measures the proportion of transactions containing both itemsets, confidence measures the conditional probability of BB given AA, and lift compares that confidence with the overall frequency of BB. High confidence alone can be misleading when the consequent is already very common. (docs.oracle.com)

Anomaly detection identifies observations that depart from a modeled norm. Such observations may warrant investigation, but an anomaly score does not itself establish fraud or error. Feature extraction constructs new attributes from existing ones; methods such as principal component analysis can provide a reduced representation for subsequent modeling. (docs.oracle.com)

The discovery workflow

A widely used process framework is CRISP-DM, the Cross-Industry Standard Process for Data Mining. It distinguishes six phases: business understanding, data understanding, data preparation, modeling, evaluation, and deployment. These phases are iterative rather than strictly sequential: an evaluation result may reveal a need to revise the problem definition, obtain additional data, or change the model. (ibm.com)

Problem definition establishes what the analysis is intended to accomplish and how success will be assessed. Data understanding examines the available records and their suitability. Preparation transforms those records into modeling inputs. Modeling creates candidate representations or predictors, while evaluation considers whether their results address the original objectives. Deployment may involve producing a report or integrating model scores into an operational system; it does not necessarily mean automating a decision. (ibm.com)

Preprocessing is part of the analytical method, not merely an administrative step. Operations such as feature selection, missing-data imputation, and scaling can affect the final model and its apparent performance. In predictive evaluation, transformations learned from data must be fitted using only the relevant training data, then applied consistently to held-out observations. (scikit-learn.org)

Evaluation and interpretive limits

A model’s fit to observed records is not sufficient evidence of predictive usefulness. Overfitting occurs when a model captures details that do not transfer effectively to unseen cases. A separate test set and cross-validation help assess generalization, although the evaluation design must reflect the data’s structure. Time-ordered observations and repeated records from the same entity may require different splitting strategies from independent examples. (scikit-learn.org)

Data leakage occurs when model development uses information unavailable at prediction time. It can arise through target-related variables or through preprocessing and feature selection performed before separating training and evaluation data. Leakage can produce optimistic scores even when cross-validation is used. (scikit-learn.org)

Association rules express co-occurrence, not causation. A strong purchasing relationship therefore does not establish that buying one product causes another purchase. Interpretation must distinguish an observed statistical relationship from a causal explanation. (docs.oracle.com)

Privacy and data governance

Mining personal records raises questions about disclosure and the conditions under which data are used or shared. De-identification can reduce privacy risks, but it is not an absolute guarantee: research has demonstrated that some de-identified datasets can be re-identified. Governance therefore concerns both the usefulness of analytical outputs and the possibility of exposing information about individuals. NIST’s guidance treats de-identification as a combination of technical methods and organizational controls, with explicit attention to disclosure risk and the data lifecycle. (csrc.nist.gov)