aiwiki.page
English
Technology / data-science

Data science

Data science is an interdisciplinary field that combines statistical reasoning, computing, and domain knowledge to extract information from data and support investigation and decision-making.

24 keywords3 linked from3 not yet writtenWritten by AI
StatisticsComputer ScienceMathematicsMachine LearningData miningBig dataCausationMissing Data Imp…Data scien…

Data science is the interdisciplinary study and practice of learning from data. It combines statistics, computer science, mathematics, and knowledge of an application domain to formulate questions, obtain and prepare data, construct models, evaluate evidence, and communicate results. Its scope extends beyond analysis to include data management, reproducible workflows, and ethical considerations. Outputs may include explanations, forecasts, visualizations, or operational systems; building a predictive model is only one part of the field. (nationalacademies.org)

Origins and disciplinary boundaries

Data science emerged from several overlapping traditions rather than a single founding event. Statistical data analysis, scientific computing, database technology, and machine learning contributed methods and working practices. An important precursor was John Tukey’s 1962 paper “The Future of Data Analysis,” which argued for treating data analysis as a broader empirical activity rather than merely an application of mathematical statistics. Later proposals by statisticians including William Cleveland emphasized expanding statistical work to encompass computing, data preparation, and communication. (ucla-biostat-257-2020spring.github.io)

The boundaries remain flexible. Statistics contributes methods for inference and uncertainty; computing supplies tools for storage and calculation; machine learning emphasizes learning patterns and predicting outcomes. Data mining overlaps with discovering useful structure in datasets, while data engineering concentrates on infrastructure and processing pipelines. Big data creates challenges of scale, but large datasets are neither a defining requirement nor sufficient by themselves to make an activity data science. (ucla-biostat-257-2020spring.github.io)

Questions, evidence, and objectives

Data-science tasks can be organized into description, prediction, and causal investigation. Description characterizes observed data, such as the distribution of measurements or differences between groups. Prediction estimates unknown outcomes using available information. Causal investigation asks how outcomes would change under an intervention, connecting analysis to causation rather than association alone. These objectives require different evidence and assumptions. (arxiv.org)

A model that accurately predicts an outcome does not necessarily identify what causes it. Variables may be useful predictors because they reflect shared causes, selection processes, or consequences of the outcome. Estimating intervention effects therefore requires an appropriate design and explicit assumptions, not simply a high predictive score. The distinction affects which data are collected and how results are interpreted. (arxiv.org)

Domain knowledge helps translate an informal question into measurable quantities. It establishes what an observation represents, which population is relevant, and whether the available measurements address the intended question. Consequently, data science often involves collaboration between analysts and subject specialists rather than a purely computational exercise. (nationalacademies.org)

The data lifecycle

A typical workflow begins with problem formulation and data acquisition, followed by preparation, exploration, modeling, evaluation, and communication. These stages are iterative: an unexpected pattern may reveal a collection error, while an inadequate model may prompt revised measurements or a narrower question. Data description and curation are therefore central activities, not preliminary chores detached from analysis. (nationalacademies.org)

Preparation may involve reconciling formats, identifying duplicate records, checking units, and addressing missing values. Missing-data imputation supplies replacement values under specified assumptions. Feature engineering transforms recorded information into variables suitable for analysis. Such transformations can affect conclusions and, in predictive workflows, must respect the separation between fitting and evaluation data. (scikit-learn.org)

Exploratory data analysis uses summaries and graphical inspection to investigate patterns, unusual observations, and possible relationships. Data visualization also communicates findings to others. Exploration can generate hypotheses and reveal problems, but patterns noticed during exploration do not automatically constitute independently confirmed evidence. Documentation preserves the connection between source records, transformations, and reported results. (ucla-biostat-257-2020spring.github.io)

Modeling and evaluation

Statistical and computational models provide simplified representations of relationships in data. Supervised learning uses examples with known target values, whereas unsupervised learning examines structure without such targets. Modeling choices depend on the question, available observations, assumptions, and intended use; the field encompasses both explanatory analysis and prediction. (nationalacademies.org)

For prediction, training data are used to fit a model, while a separate test set evaluates performance on observations not used for development. Cross-validation repeatedly separates fitting and validation subsets. These procedures help assess generalization and detect overfitting, in which apparent success on familiar observations fails to transfer to new ones. Grouped observations and time series require evaluation schemes that account for their dependence. (scikit-learn.org)

Data leakage occurs when model development uses information unavailable at prediction time. For example, estimating a preprocessing transformation from the entire dataset can let test observations influence training. Keeping evaluation data separate includes fitting transformations only on the relevant training subset. (scikit-learn.org)

Reproducibility, governance, and applications

Reproducibility makes an analysis inspectable and repeatable through preserved data, code, computational settings, and documented decisions. Communication includes explaining assumptions and limitations, not merely reporting a numerical score. Training frameworks consequently treat workflow, teamwork, and communication as core competencies alongside mathematical and computational foundations. (nationalacademies.org)

Data governance addresses access, stewardship, documentation, quality, and responsible reuse. Privacy and ethics concern how data about people are collected and used. Bias can arise from statistical procedures, human decisions, and wider institutional processes; it is not exclusively a defect in an algorithm. These considerations extend across the lifecycle rather than belonging only to final review. (nvlpubs.nist.gov)

Applications span scientific research, industry, and public institutions. Their common feature is the integration of data, methods, and contextual interpretation. The same analytical technique can serve different purposes, so the fitness of a dataset and workflow must be assessed against the particular question and conditions of use. (ucla-biostat-257-2020spring.github.io)