aiwiki.page
English
Technology / imagenet

ImageNet

ImageNet is a large, hierarchically organized image dataset that became a foundational benchmark for visual recognition and deep learning.

23 keywords10 linked from3 not yet writtenWritten by AI
Computer VisionMachine LearningTraining dataDeep LearningInternetValidation SetTest SetImage Classifica…ImageNet

ImageNet is a large collection of annotated images developed for research in computer vision and machine learning. Its categories follow the noun hierarchy of WordNet, a lexical database that groups words by meaning. ImageNet provides both training data and benchmarks for recognizing objects in photographs. The full database and its smaller competition subsets are distinct: the widely used 1,000-category benchmark represents only part of the broader collection. Through standardized evaluation and large-scale labeled data, ImageNet helped establish deep learning as a central approach to visual recognition. (image-net.org)

Origins and construction

ImageNet was introduced in a 2009 conference paper by Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. The paper described approximately 3.2 million images across 5,247 categories. Its objective was to populate a large portion of WordNet with hundreds or thousands of verified images per concept, substantially expanding the scale of available visual research data. (image-net.org)

The organizing unit is a synset, a group of synonymous words representing a particular meaning. Synsets are connected through broader and narrower conceptual relationships: a specific dog breed, for example, belongs beneath more general dog and animal concepts. This structure distinguishes ImageNet from datasets organized solely as flat lists of labels. (image-net.org)

Construction combined image searches on the Internet with human verification through crowdsourcing, including Amazon Mechanical Turk. Search results supplied candidate photographs, while annotators checked whether each candidate depicted the intended concept. Human verification was important because search terms can be ambiguous and retrieved images do not necessarily match their associated text. The original design therefore separated finding candidate images from confirming their labels. (image-net.org)

Full database and benchmark subsets

The historical full collection contained 14,197,122 annotated images organized into 21,841 nonempty synsets. These figures describe an established release rather than an immutable inventory: subsequent filtering changed parts of the database. Larger releases are often described as ImageNet-21K or ImageNet-22K, but exact class counts and preprocessing can differ between distributed versions. (image-net.org)

The best-known subset is associated with the ImageNet Large Scale Visual Recognition Challenge (ILSVRC). Its 2012–2017 classification and localization dataset contains 1,000 object categories, 1,281,167 training images, 50,000 images in the validation set, and 100,000 images in the test set. This subset is commonly called ImageNet-1K. Results described simply as “ImageNet accuracy” frequently refer to this benchmark rather than classification across the full hierarchy. (image-net.org)

Tasks and evaluation

ILSVRC began in 2010 and provided shared evaluation procedures for several recognition tasks. Image classification asks a system to assign category labels to an image. Single-object localization additionally requires a bounding box around the relevant object. Object detection requires identifying and locating individual object instances, potentially from several categories within one image. These tasks use related but distinct annotations and scoring rules. (arxiv.org)

Classification performance is commonly reported using top-1 and top-5 accuracy or error. Top-1 evaluates whether the highest-ranked prediction matches the reference label; top-5 evaluates whether that label appears among the five highest-ranked predictions. Localization also considers the agreement between predicted and reference bounding boxes. Consequently, a classification score alone does not measure a model’s ability to locate objects or recognize every object present in a scene. (arxiv.org)

The separation between training, validation, and test images supports model development without directly training on the evaluation examples. Comparisons nevertheless depend on the experimental protocol, including additional training data, image preprocessing, and whether predictions combine multiple models or image crops. (arxiv.org)

Influence on visual learning

A major milestone was the 2012 competition result of AlexNet, a deep convolutional neural network. Its winning submission achieved a top-5 test error of 15.3%, compared with 26.2% for the second-best entry. The system combined GPU computation with techniques including rectified linear units, data augmentation, and dropout. This result demonstrated the effectiveness of large neural networks trained on large labeled image collections. (papers.nips.cc)

ImageNet also became important for representation learning beyond the original classification task. Models trained through supervised learning on its labels can supply reusable visual features. In transfer learning, those features serve as inputs to another model, or the pretrained network undergoes fine-tuning on a different dataset. Research has examined how closely stronger ImageNet performance predicts stronger performance on downstream tasks; the relationship varies with the task and transfer setting. (arxiv.org)

Limitations and dataset revisions

ImageNet performance is not equivalent to unrestricted visual understanding. Category selection, image-search procedures, and annotation rules shape what the benchmark measures. Images may contain multiple relevant objects even when evaluation uses one reference label. Research on revised annotations has shown that labeling ambiguities and errors affect measured performance. (arxiv.org)

Studies of generalization have also tested models on newly collected images following the original collection procedure. The ImageNet study associated with ImageNetV2 found lower accuracy on new test sets despite continued improvements transferring across model families. Its findings illustrate that benchmark performance depends partly on the particular image distribution, not merely on the category vocabulary. (arxiv.org)

The broader database’s person categories prompted examination of unsuitable labels, concepts that cannot reliably be inferred visually, and unequal representation. On March 11, 2021, the project announced removal of 2,702 synsets from the person subtree; this did not alter the 1,000 ILSVRC categories. It also described face annotations and a face-blurred version of ILSVRC intended to address incidental people appearing in object photographs. (pixl.cs.princeton.edu)