Computer vision is a field of computer science concerned with extracting meaningful information from images, video, and other visual measurements. Its methods identify objects, estimate motion and three-dimensional structure, and interpret relationships within scenes. Closely connected to artificial intelligence, it combines computational models of image formation with statistical and learning-based approaches. Outputs may be category labels, object locations, measurements, or reconstructed environments rather than images themselves. (link.springer.com)
Foundations and scope
A digital image records measurements on a grid of pixels, commonly with separate color channels. These measurements depend on illumination, surface properties, camera position, and sensor characteristics. Recovering information about the world therefore requires accounting for how observations were formed. Geometry describes spatial relationships and projection, while optics describes aspects of light and imaging. Many vision tasks are inverse problems: different physical scenes can produce similar observations, making additional assumptions or measurements necessary. (cs.jhu.edu)
Computer vision overlaps with image processing, which includes filtering, enhancement, and image transformation, but emphasizes inference about depicted content. It also overlaps with computer graphics, particularly in reconstruction and rendering. The boundaries are not absolute: image restoration can support recognition, and geometric reconstruction can supply models for graphics. (cs.jhu.edu)
Principal tasks
Image classification assigns an image to one or more categories, such as identifying its dominant object or scene type. Object detection additionally locates individual objects, commonly using bounding boxes. A photograph containing several vehicles may consequently receive one image-level label but several separate detections. (cs231n.stanford.edu)
Semantic segmentation assigns a category to each pixel, distinguishing regions such as roads, buildings, and vegetation. Instance segmentation also separates individual objects belonging to the same category. Pose estimation locates meaningful landmarks, such as joints on a person, while tracking associates objects across successive video frames. These tasks differ in the structure and detail of their required outputs. (cs231n.stanford.edu)
Geometric tasks include estimating depth, recovering camera motion, and reconstructing three-dimensional scenes. Stereo vision uses differences between corresponding observations from different viewpoints. Optical flow estimates apparent image motion between frames; this motion can reflect both moving objects and camera movement. Image registration aligns observations, supporting applications such as panoramic stitching and comparisons over time. (cs.jhu.edu)
Methods and historical development
Traditional pipelines often combine preprocessing, feature extraction, matching, and a task-specific decision rule. Features describe edges, corners, textures, or distinctive local patterns. Geometric methods can then estimate transformations or scene structure, while classifiers use these descriptors to recognize categories. Such approaches explicitly separate the construction of visual representations from subsequent inference. (cs.jhu.edu)
Machine learning introduced data-driven models that learn decision rules from examples. Deep learning extends this approach by learning multilayer representations alongside the final prediction. Convolutional neural networks exploit local image structure and reuse filters across spatial positions, reducing the number of parameters compared with similarly arranged fully connected networks. (papers.nips.cc)
In 2012, AlexNet demonstrated substantially improved large-scale recognition performance in the ImageNet Large Scale Visual Recognition Challenge. Its training combined a deep convolutional architecture, graphics processing units, and techniques for limiting overfitting. The associated ImageNet dataset provided large quantities of labeled images for learning and evaluation. (papers.nips.cc)
The Vision Transformer, introduced in a 2020 paper, showed that image recognition could also use a transformer architecture operating on sequences of image patches. Its results demonstrated the effectiveness of large-scale pretraining followed by transfer to downstream recognition tasks, without requiring a convolutional backbone. (arxiv.org)
Training and evaluation
In supervised learning, examples pair images with annotations appropriate to the task: category labels, bounding boxes, pixel masks, or landmarks. A loss function measures disagreement between predictions and targets. Neural-network training adjusts parameters through backpropagation and numerical optimization. Annotation requirements differ substantially across tasks, affecting dataset construction and training cost. (cs231n.github.io)
Data augmentation creates variations of training examples, for instance through cropping, reflection, or color modification. These transformations must preserve the intended prediction target. Transfer learning reuses representations learned on another dataset, often adapting a pretrained network rather than learning all parameters from scratch. Both techniques can be useful when labeled examples are limited. (papers.nips.cc)
A validation set supports model selection, while a separate test set estimates performance after those choices are fixed. Classification accuracy measures the proportion of correct predictions; detection evaluation also considers localization and precision–recall behavior. Segmentation commonly measures overlap between predicted and reference regions. Scores are meaningful only in relation to the dataset, annotation rules, and evaluation protocol. (cs231n.github.io)
Applications and limitations
Applications include manufacturing inspection, document analysis, photographic stitching, visual search, and three-dimensional modeling. In robotics, visual measurements help systems estimate position and perceive their surroundings. Different applications impose different requirements for accuracy, processing speed, and spatial precision. (link.springer.com)
Recognition remains sensitive to viewpoint, illumination, occlusion, background clutter, and variation within categories. A model must distinguish relevant differences while tolerating changes that should not alter its prediction. Strong performance on one dataset consequently does not establish equivalent performance under all operating conditions. (cs231n.github.io)
Face-recognition evaluations illustrate the importance of application-specific testing. NIST’s 2019 study documented demographic differences in error rates across many tested algorithms, with results varying by algorithm and matching task. NIST also identifies image quality as an important influence on false-negative rates. Such findings concern particular systems and evaluation conditions rather than a uniform property of every vision method. (nist.gov)