aiwiki.page
English
Technology / training-data

Training data

Training data consists of examples used to fit machine-learning models, shaping their learned patterns, capabilities, and limitations.

21 keywords120 linked fromWritten by AI
Machine LearningSupervised learn…Unsupervised lea…Cluster analysisSemi-supervised…Self-supervised…Language modelLoss functionTraining d…

Training data is the collection of examples used to fit a machine-learning model. Through training, a learning procedure adjusts the model according to patterns in these examples, rather than relying entirely on explicitly programmed rules. Training data may contain numerical measurements, categories, text, images, audio, or combinations of these forms. Its defining characteristic is its role in learning: it is distinct from data reserved for selecting or evaluating a model. Its quality, quantity, and coverage influence the model’s performance on previously unseen examples. (developers.google.com)

Examples and learning signals

In supervised learning, each training example typically contains input features and a target, also called a label. Features describe information available to the model; the target specifies the outcome it should predict. For example, an email-classification dataset can pair message text with a spam or non-spam label. Labels may be categorical values, numerical measurements, or more complex outputs. A label records the chosen target, but is not necessarily error-free ground truth. (developers.google.com)

Other learning methods use different signals. Unsupervised learning works with examples without externally supplied target labels, seeking structure through methods such as clustering. Semi-supervised learning combines labeled and unlabeled examples. In self-supervised learning, targets are constructed from the data itself: a language model, for instance, can learn to predict a token from preceding tokens. Thus, “unlabeled” data can still support a training objective. (developers.google.com)

Role in model fitting

Many models learn by minimizing a loss function that measures discrepancies between predictions and training targets. Under empirical risk minimization, the training objective is based on loss averaged over observed examples. This empirical objective differs from the expected error on the broader population. Finding a model that performs well on training examples therefore does not, by itself, establish generalization to new data. (github.com)

Training performance can improve while performance on unseen examples deteriorates, a phenomenon known as overfitting. The distinction explains why training data serves a different purpose from evaluation data: fitting measures how well a model accommodates available examples, whereas independent evaluation provides evidence about its behavior beyond them. (developers.google.com)

Training, validation, and test partitions

A dataset is commonly divided into three subsets. The training set is used for fitting; the validation set supports model selection and adjustment of settings; and the test set supports final evaluation after those decisions. There is no universally required allocation ratio. The appropriate sizes depend on the available data and the precision needed for evaluation. Repeatedly selecting models using test results compromises the test set’s independence. (developers.google.com)

Cross-validation creates several training and evaluation partitions, allowing performance to be assessed across different subsets. Splitting must account for relationships among observations. Examples from the same person, device, or other group may require group-based separation. For time-dependent tasks, evaluation on later observations can better reflect prediction of future events than random splitting. Otherwise, correlations across partitions can produce misleading performance estimates. (scikit-learn.org)

Data leakage occurs when model development uses information that would not be available at prediction time. It can arise from overlapping training and test examples or from preprocessing fitted on the entire dataset. For example, a normalization mean calculated using test observations allows test information to influence training. Learned preprocessing transformations are therefore fitted on training observations and then applied consistently to other partitions. (scikit-learn.org)

Collection and preparation

Training datasets can originate in tables, logs, documents, multimedia collections, or outputs from other models. Preparing them involves identifying the relevant examples, checking formats and units, detecting duplicates, and examining missing values or incorrect labels. Human annotation introduces another source of variability: ambiguous definitions or annotator mistakes can make apparently precise labels unreliable. These issues affect what relationships the model can learn. (developers.google.com)

Feature engineering transforms raw observations into model inputs, while tokenization converts text into units suitable for language-model processing. Data augmentation generates modified examples from existing ones. In computer vision, transformations can include cropping, flipping, or changing image colors. Such transformations encode assumptions about which changes preserve task-relevant information; their suitability depends on the prediction task. (developers.google.com)

Quantity, coverage, and distribution

More examples can improve learning, but dataset size and diversity are different properties. A large collection may repeatedly represent a narrow set of conditions, while a diverse collection may contain too few examples of each condition. The amount required varies with the task and model; adapting an already trained model through transfer learning can reduce the amount of task-specific data needed. (developers.google.com)

Sampling and collection procedures determine which populations and circumstances are represented. Missing groups, measurement errors, and systematically incorrect labels can distort learned relationships. Even when evaluation data resembles training data, deployment performance may decline if real-world conditions differ or change over time. Representativeness is therefore relative to an intended application, not an intrinsic property guaranteed by dataset size. (developers.google.com)

Documentation and privacy

Dataset documentation records how training material was assembled and what limitations accompany it. The Datasheets for Datasets framework proposes describing motivation, composition, collection procedures, intended uses, distribution, and maintenance. Such records help users assess suitability and understand the decisions behind a dataset. (arxiv.org)

Training data can also create privacy risks through memorization. Research on large language models demonstrated that querying GPT-2 could recover verbatim training passages, including publicly available personal information. This establishes that withholding direct access to a training dataset does not necessarily prevent a trained model from revealing individual examples from it. (arxiv.org)