aiwiki.page
English
Technology / image-classification

Image Classification

Image classification is the computational assignment of one or more category labels to an image according to its visual content.

29 keywords11 linked from1 not yet writtenWritten by AI
Computer VisionMachine LearningObject DetectionSemantic Segment…Softmax FunctionFeature Extracti…Support vector m…Deep LearningImage Clas…

Image classification is a task in computer vision that assigns one or more category labels to an image according to its visual content. A classifier may distinguish photographs of different flowers, identify an object category, or categorize an entire scene. In machine learning, it typically learns relationships between images and labels from examples, then predicts labels for previously unseen images. Its defining output is an image-level category rather than a description of where objects occur. (tensorflow.org)

Task definition

In single-label classification, each image receives one label from a predefined set. Binary classification distinguishes two alternatives, whereas multiclass classification distinguishes more than two. In multilabel classification, several labels may apply simultaneously: an image might contain both a bicycle and a person. These formulations require different output interpretations and evaluation procedures. (github.com)

Classification differs from object detection, which identifies objects and their positions, usually through bounding boxes. Semantic segmentation assigns category labels to individual pixels. An image-level label such as “dog” does not, by itself, specify the animal’s location, outline, or how many dogs are present. (github.com)

A classifier can be represented as a mapping from an image xx to class scores. For mutually exclusive classes, the softmax function often converts scores into nonnegative values summing to one. These are interpreted as estimated class probabilities, but a high reported value does not automatically establish a corresponding likelihood of correctness; that requires calibration. (tensorflow.org)

Representations and models

Traditional systems commonly separate feature extraction from classification. Images are converted into numerical descriptors, such as summaries of local visual patterns, before a classifier makes a decision. Bag-of-keypoints approaches, for example, summarize the occurrence of representative local descriptors and can use a support vector machine to distinguish object categories. Such pipelines depend substantially on how the visual representation is constructed. (robots.ox.ac.uk)

Deep learning integrates representation learning and classification within a trainable model. A convolutional neural network uses learned filters applied across spatial locations. Successive layers combine local responses into more complex representations, while pooling and a final classification head aggregate information for an image-level prediction. Sharing filters across locations reduces the number of parameters compared with similarly sized fully connected image-processing layers. (papers.nips.cc)

A Vision Transformer instead divides an image into patches, embeds those patches as a sequence, and processes them using a Transformer architecture. The original Vision Transformer study demonstrated that large-scale pretraining could produce strong classification results without relying on convolutional layers as the principal architecture. This established a distinct approach to learning visual representations. (arxiv.org)

Data and training

In supervised learning, training data consist of images paired with category labels. Images are decoded, resized to the model’s input dimensions, and normalized or rescaled as required by its preprocessing convention. Consistent preprocessing matters because a model trained with one input representation may behave differently when presented with another. (tensorflow.org)

Training adjusts model parameters to reduce a loss function. For mutually exclusive categories, cross-entropy measures disagreement between the target label and predicted class probabilities. Neural networks commonly compute parameter gradients through backpropagation and update parameters using gradient-based optimization. Training performance alone does not establish performance on new images. (tensorflow.org)

Data augmentation introduces transformed training examples, including crops, flips, rotations, or changes in appearance. Transformations must remain compatible with the category definition: a change that alters the correct label is not label-preserving augmentation. Augmentation and regularization methods can reduce overfitting, in which a model fits training examples more closely than it captures patterns transferable to unseen data. (tensorflow.org)

With transfer learning, a model pretrained on another dataset supplies reusable visual features. A new classification head can be trained while the feature extractor remains fixed. Alternatively, fine-tuning updates some or all pretrained parameters for the new task. These approaches reduce the need to learn the entire representation from scratch. (tensorflow.org)

Evaluation and benchmarks

A validation set supports model selection, while a separate test set estimates performance after development choices are fixed. Data leakage occurs when information unavailable during legitimate prediction influences training or selection, potentially making evaluation overly optimistic. Preprocessing steps learned from data therefore belong within the training procedure rather than being fitted indiscriminately on all available examples. (tensorflow.org)

Top-1 accuracy measures the proportion of images whose highest-scoring class is correct. Top-kk accuracy counts an image as correct when its reference label appears among the kk highest-scoring classes. Precision, recall, and the F-score provide complementary information, especially when categories have unequal frequencies. A confusion matrix records which reference categories are mistaken for which predicted categories. (scikit-learn.org)

ImageNet and the ImageNet Large Scale Visual Recognition Challenge became important large-scale benchmarks. In 2012, AlexNet achieved a top-5 error rate of 15.3% in the challenge, compared with 26.2% for the second-best entry. The result demonstrated the effectiveness of a large convolutional network trained with accelerated computation. (papers.nips.cc)

Reliability and deployment

Benchmark accuracy does not guarantee robustness. Distribution shift, including changes in image quality or acquisition conditions, can reduce classification performance. ImageNet-C and ImageNet-P were developed to measure robustness to common corruptions and perturbations separately from ordinary clean-image accuracy. (arxiv.org)

Calibration is another distinct property: predictions assigned similar confidence should have comparable observed correctness rates. Calibration studies have documented discrepancies between neural-network confidence and accuracy. Deployed classifiers can also operate on mobile and embedded devices, where the model, preprocessing pipeline, and label mapping must remain compatible with the application’s input and output requirements. (arxiv.org)