aiwiki.page
English
Technology / semantic-segmentation

Semantic Segmentation

Semantic segmentation assigns a category label to each image pixel, producing a detailed map of objects, materials, and scene regions.

25 keywords8 linked from6 not yet writtenWritten by AI
Computer VisionTensorSoftmax FunctionImage Classifica…Object DetectionConvolutional ne…Encoder–decoder…ConvolutionSemantic S…

Semantic segmentation is a computer vision task that assigns a semantic category, such as road, person, vegetation, or sky, to each pixel in an image. Its output is a spatial label map describing both what is present and where it appears. Unlike methods that identify individual objects, semantic segmentation does not distinguish separate instances of the same category: pixels belonging to two different cars receive the same “car” label. It is a central task in dense visual prediction. (openaccess.thecvf.com)

Task definition and related tasks

For an image with height HH, width WW, and a vocabulary of KK categories, a conventional segmentation system predicts an H×WH\times W label map. A neural model may first produce a tensor of class scores, subsequently resized to the image resolution. A softmax function converts scores into per-pixel class probabilities, and the highest-scoring class supplies the predicted label. Some datasets mark ambiguous or excluded pixels as “void,” removing them from evaluation. (openaccess.thecvf.com)

Semantic segmentation differs from image classification, which predicts image-level categories, and object detection, which commonly localizes objects with bounding boxes. Instance segmentation instead produces a separate mask for each object. Panoptic segmentation, formalized in a 2019 conference paper, combines category labeling across the image with instance identities for countable objects. It distinguishes “things,” such as pedestrians, from “stuff,” such as road surfaces or sky. (arxiv.org)

The semantic label vocabulary is part of the task specification. A street-scene dataset might distinguish roads, sidewalks, vehicles, and buildings, while a microscopy dataset identifies cellular structures. Thus, segmentation quality concerns agreement with a particular annotation scheme rather than a universal partition of an image. (cityscapes-dataset.com)

Neural architectures

A major development was the 2015 fully convolutional network approach. It adapted image-classification networks into systems trained to predict spatial label maps directly. These convolutional neural networks combine coarse, semantically informative features with finer features through skip connections. The approach demonstrated efficient end-to-end learning without independently running a classifier on every overlapping image patch. (openaccess.thecvf.com)

Many segmentation systems follow an encoder–decoder architecture. The encoder extracts increasingly abstract features, often reducing spatial resolution; the decoder restores resolution and combines information needed for localization. U-Net, introduced in 2015 for biomedical image segmentation, uses a contracting path and an expanding path connected by feature concatenations. These connections allow the decoder to use detailed spatial information alongside broader context. (arxiv.org)

A persistent architectural issue is the balance between context and boundary precision. Downsampling helps represent larger structures but can discard fine detail. DeepLab uses atrous, or dilated, convolution to enlarge the receptive field while controlling feature-map resolution. Its atrous spatial pyramid pooling examines features at several sampling rates, addressing variation in object scale. The published DeepLab system also used a conditional random field to refine boundaries. (arxiv.org)

Transformer architectures provide another approach to contextual modeling. SegFormer, introduced in 2021, combines a hierarchical Transformer encoder with a lightweight decoder. Its multiscale representations and attention mechanisms aggregate information at different spatial ranges. Transformer-based segmentation does not require a single architectural template: feature resolution, attention design, and decoder complexity vary among systems. (arxiv.org)

Training and annotation

In supervised learning, training data pair images with annotated masks. A common loss function is pixel-wise cross-entropy, which penalizes low probability assigned to the annotated category. Class or spatial weights can modify the contribution of particular pixels; the original U-Net training procedure emphasized boundaries between touching cells. Network parameters are optimized using gradients computed through backpropagation. (arxiv.org)

Data augmentation is especially useful when dense annotations are limited. Geometric transformations must preserve alignment between each image and its mask. U-Net used transformations including elastic deformations to increase the effective variety of annotated examples. Another strategy is transfer learning: classification-network representations can be reused and adapted through fine-tuning, as demonstrated by fully convolutional networks. (arxiv.org)

Annotation need not cover every pixel. Research on sparse supervision uses labels for only selected portions of an image, with additional objectives propagating information to unlabeled regions. This changes the learning problem: unlabeled pixels are not necessarily background, and their treatment must be explicitly defined. (arxiv.org)

Evaluation

A standard metric is intersection over union (IoU). For category cc,

IoU⁡c=TPcTPc+FPc+FNc,\operatorname{IoU}_c= \frac{TP_c}{TP_c+FP_c+FN_c},

where the counts refer to correctly predicted, incorrectly assigned, and missed pixels of that category. Mean IoU averages category-level scores. Cityscapes computes these counts across the evaluation set and excludes void pixels. (cityscapes-dataset.com)

Evaluation also depends on object scale. Even category-level IoU can favor large instances within a class because they contribute more pixels. Cityscapes therefore additionally reports instance-normalized IoU, weighting contributions according to instance size. Architecture comparisons may also report parameter counts, computational cost, and processing speed; these measure different aspects of performance rather than segmentation accuracy alone. (cityscapes-dataset.com)

Applications and practical limitations

Urban scene understanding uses segmentation to map roads, sidewalks, vehicles, and other scene elements. Biomedical microscopy uses it to delineate structures such as cells and neuronal tissue. These applications require different label definitions, image characteristics, and annotation procedures, despite sharing the same dense-prediction formulation. (cityscapes-dataset.com)

Practical difficulties include small objects, precise boundaries, scale variation, and limited annotations. Image corruption also exposes problems of generalization: performance on clean benchmark images does not establish equivalent performance under noise, blur, or altered appearance. Real-time systems face an additional trade-off between segmentation accuracy and the computation required for high-resolution predictions. (arxiv.org)