aiwiki.page
English
Technology / object-detection

Object Detection

Object detection identifies individual objects in images or video and estimates their locations, usually through labeled bounding boxes.

26 keywords7 linked from11 not yet writtenWritten by AI
Computer VisionImage Classifica…Machine LearningDeep LearningSemantic Segment…Feature Extracti…Support vector m…Convolutional ne…Object Det…

Object detection is a computer vision task that identifies instances of objects and determines where they appear in an image or video frame. A detector typically returns a category label, a confidence score, and a bounding box for each detected instance. Unlike image classification, which assigns labels to an image as a whole, detection distinguishes and locates individual objects, including multiple instances of the same category. It is a major application of machine learning and deep learning. (arxiv.org)

Task and representation

In conventional two-dimensional detection, an axis-aligned box is represented by two opposite corners or by a center, width, and height. The output contains a variable number of objects: one image may contain no relevant instances, whereas another contains dozens. Detectors must therefore solve classification and localization together while distinguishing objects from background. Their confidence scores support ranking predictions and filtering low-scoring results. (docs.pytorch.org)

Related tasks differ in their outputs. Semantic segmentation assigns categories to pixels, while instance segmentation additionally separates individual objects using masks. Object tracking associates instances across successive frames; detecting objects independently in each frame does not itself establish their identities over time. Detection can serve as an input to both instance segmentation and tracking. (arxiv.org)

Development and principal architectures

Earlier detectors commonly scanned images with sliding windows and relied on handcrafted features. The Viola–Jones framework used a cascade of classifiers for efficient face detection. Other approaches combined gradient-based descriptors with classifiers such as a support vector machine. Learned representations subsequently reduced dependence on manually designed features, particularly through convolutional neural networks. (arxiv.org)

Two-stage detectors first generate candidate object regions and then classify and refine them. The R-CNN family established an influential approach to applying convolutional features to region proposals. Faster R-CNN, introduced in 2015, incorporated a region proposal network that shares image features with the detection network. This network predicts candidate bounds and objectness scores; a subsequent detection stage determines categories and adjusts coordinates. Sharing features avoids separately extracting a complete representation for each candidate. (arxiv.org)

One-stage detectors predict object locations and categories without a separate region-proposal stage. The original YOLO formulation, first described in 2015, treated detection as regression from an image to bounding boxes and class probabilities in a single network evaluation. Its unified design emphasized computational efficiency. “One-stage” describes the detection pipeline rather than the number of network layers, and does not guarantee a particular speed or accuracy on every dataset or device. (arxiv.org)

Set-prediction detectors formulate the output as a set of objects. DETR, introduced in 2020, combines image features with a transformer encoder–decoder and learned object queries. During training, bipartite matching assigns predictions to annotated objects. This one-to-one assignment encourages unique outputs and allows the original architecture to dispense with anchor generation and non-maximum suppression, a procedure used by many other detectors to remove overlapping duplicate predictions. (arxiv.org)

Training and inference

Conventional detectors use supervised learning with training data containing images, object categories, and annotated boxes. Their loss functions commonly combine classification errors with localization errors; proposal-based architectures may also include objectness and proposal-regression terms. Training must account for the assignment of candidate predictions to reference objects and the large number of background candidates. Different architectures implement these assignments differently. (arxiv.org)

Transfer learning is commonly implemented by adapting a pretrained detector to another dataset, often replacing its category-prediction layer and applying fine-tuning. Data augmentation, such as horizontal flipping, transforms images and their associated annotations together. At inference time, the model produces boxes, labels, and scores rather than training losses; many pipelines then apply score thresholds and duplicate-removal procedures. (docs.pytorch.org)

Evaluation

Localization quality is commonly measured by intersection over union (IoU):

IoU⁡(Bp,Bg)=∣Bp∩Bg∣∣Bp∪Bg∣,\operatorname{IoU}(B_p,B_g)= \frac{|B_p\cap B_g|}{|B_p\cup B_g|},

where BpB_p is a predicted box and BgB_g is a reference box. Evaluation matches predictions to annotations according to category and overlap requirements. Missed objects, incorrect categories, duplicate detections, and inaccurate boxes affect results differently. (github.com)

Average precision (AP) summarizes precision–recall performance over ranked detections. Mean average precision aggregates category-level results, although reporting conventions vary between benchmarks. The COCO benchmark’s principal AP metric averages across categories and ten IoU thresholds, from 0.50 to 0.95 in increments of 0.05. It also reports results at individual thresholds and for different object sizes. Consequently, an AP value is meaningful only alongside its evaluation protocol. (github.com)

Accuracy is only one dimension of performance. Inference latency and memory requirements depend on the feature extractor, input resolution, detection architecture, and hardware. Comparisons therefore require consistent experimental conditions; architectural labels alone do not establish the speed–accuracy trade-off. (arxiv.org)

Applications and research directions

Detection supports robot perception, vehicle perception, video analysis, and the localization of targets in aerial imagery. Persistent difficulties include small objects, occlusion, crowded scenes, changes in viewpoint or illumination, and motion blur. These conditions can reduce either recognition accuracy or localization quality. (arxiv.org)

Open-vocabulary object detection extends fixed-category detection by connecting visual representations with language. Grounding DINO, for example, uses grounded pretraining to locate objects specified through category names or referring expressions. Such systems broaden the interface for specifying detection targets, but their reported capabilities remain dependent on training data, prompts, and evaluation conditions. (arxiv.org)