Machine learning is a field of artificial intelligence concerned with computational methods that improve performance through data or experience. Instead of specifying every decision rule explicitly, developers provide learning procedures, examples, and objectives from which a system constructs a model. Such models can predict outcomes, identify structure, generate content, or guide actions. The field draws on statistics and computer science; learning is assessed relative to a particular task, experience, and measure of performance, rather than assumed to imply human-like understanding. (cs.cmu.edu)
Historical development
Early research explored whether computers could improve their behavior rather than merely execute fixed instructions. In 1959, Arthur Samuel published “Some Studies in Machine Learning Using the Game of Checkers,” describing programs that improved their play through experience. His work became an early demonstration of machine learning in a constrained environment. Subsequent research developed methods for learning rules, decision trees, probabilistic relationships, and neural-network parameters. (ieeexplore.ieee.org)
During the following decades, statistical methods and neural approaches developed alongside one another. Increased computing capacity, larger datasets, and improved training methods supported the expansion of deep learning, which uses multiple processing layers to learn increasingly abstract representations. A 2015 review documented major advances in image recognition, speech recognition, and other applications, connecting these achievements to multilayer representation learning and effective optimization. (nature.com)
Learning paradigms
In supervised learning, examples contain inputs and associated targets, often called labels. A model learns a relationship that can be applied to new inputs. Classification predicts categories, such as whether a message is spam; regression predicts numerical quantities, such as an estimated price. Labels specify the learning target, but their presence does not guarantee that they are accurate or representative. (developers.google.com)
Unsupervised learning investigates data without supplied target labels. It can identify groups of similar observations through cluster analysis, compress representations, or characterize patterns in a dataset. The resulting structure depends on the chosen method and assumptions: a mathematical grouping does not automatically establish a meaningful real-world category. (developers.google.com)
In reinforcement learning, an agent selects actions in an environment and learns from rewards or penalties. Its objective concerns accumulated reward rather than simply matching labeled examples. Decisions may affect later observations and opportunities, making sequential interaction central to the problem. Generative systems, meanwhile, learn patterns that support producing new content; their training can combine different learning paradigms. (developers.google.com)
Models and training
A model is a mathematical or computational representation learned from training data. Model families include linear regression, decision trees, probabilistic classifiers, instance-based methods, and artificial neural networks. They differ in expressive capacity, assumptions, computational requirements, and how they represent relationships. Neural networks are therefore one family within machine learning, not a synonym for the entire field. (cs.cmu.edu)
Training frequently involves minimizing a loss function that measures disagreement between predictions and targets. Through mathematical optimization, a learning procedure adjusts model parameters to reduce this objective. In neural networks, backpropagation computes gradients that support parameter updates. Other methods may construct trees, estimate probabilities, or retain examples rather than optimize all parameters through gradients. (cs.cmu.edu)
Input representation also matters. Feature engineering constructs useful variables from raw observations, while representation learning enables models to learn useful transformations themselves. Deep networks can learn several levels of representation, such as combining simple visual features into more complex patterns. Neither approach removes dependence on the information available in the original data. (nature.com)
Generalization and evaluation
The central evaluation problem is generalization: whether a learned model performs usefully on examples beyond those used to fit it. Overfitting occurs when a model fits training-specific details that do not transfer reliably. Regularization constrains or penalizes aspects of the model to influence this behavior, but successful training alone cannot establish performance on new data. (cs.cmu.edu)
A typical experiment separates data into training, validation, and test sets. Validation supports model selection and configuration choices; testing estimates performance after those choices. Cross-validation repeats fitting and evaluation across partitions. For time-dependent or grouped observations, partitions must account for temporal order or group relationships rather than assume that any random split is appropriate. (scikit-learn.org)
Data leakage occurs when model development uses information unavailable at prediction time. It can arise when preprocessing is fitted on the entire dataset or when test results influence model selection. Such leakage produces misleadingly favorable estimates. Evaluation metrics must also match the intended task and conditions: a single aggregate score may conceal important differences between data segments or error types. (scikit-learn.org)
Applications and limitations
Machine learning supports computer vision, speech recognition, translation, recommendation, robotic control, and content generation. These applications differ in their objectives, available feedback, and consequences of error. A system effective in one setting may fail under different operating conditions, so deployment performance requires assessment beyond training results. (nature.com)
Reliability depends on data quality, realistic evaluation, and the relationship between training conditions and actual use. Learned systems can reproduce harmful biases, expose privacy risks, or behave unreliably in unexpected circumstances. Ongoing monitoring examines whether deployed performance remains consistent with the intended application. The NIST AI Risk Management Framework treats validity, reliability, security, transparency, explainability, privacy, and management of harmful bias as distinct considerations; predictive accuracy alone does not establish all of them. (nvlpubs.nist.gov)