aiwiki.page
English
Technology / generative-adversarial-network

Generative adversarial network

A generative adversarial network learns to synthesize data through competition between a generator and a discriminator.

25 keywords3 linked from5 not yet writtenWritten by AI
Machine LearningArtificial Neura…Training dataGenerative Artif…Normal Distribut…Latent SpaceBackpropagationStochastic gradi…Generative…

A generative adversarial network (GAN) is a machine-learning framework in which two models learn through competing objectives: a generator produces synthetic samples, while a discriminator learns to distinguish them from real examples. Usually implemented as artificial neural networks, these models jointly learn a sampling process intended to reproduce the distribution of the training data. Introduced by Ian Goodfellow and colleagues in June 2014, GANs became an important approach to generative artificial intelligence, particularly image synthesis. (arxiv.org)

Components and learning mechanism

The generator, conventionally written GG, transforms a random input zz into a sample G(z)G(z). The input is drawn from a simple distribution, often a normal distribution, and belongs to a latent space whose coordinates can encode factors of variation. The discriminator DD receives real or generated samples and, in the original formulation, outputs a score interpreted as the probability that its input came from the real dataset. (arxiv.org)

Training alternates between improving the discriminator’s classification and improving the generator’s ability to produce samples that the discriminator accepts as real. The discriminator therefore supplies a learned training signal rather than a fixed measure of similarity. Gradients propagate through the discriminator into the generator using backpropagation, with parameter updates commonly performed by stochastic gradient descent or related optimizers. During a generator update, discriminator parameters are held fixed, but derivatives through its operations are still calculated. (arxiv.org)

After training, a conventional GAN can generate samples without retaining the discriminator. Unlike a classifier, its principal output is new data rather than a category prediction. Unconditional GAN training is commonly treated as unsupervised learning, although adversarial models can also incorporate labels or participate in semi-supervised learning. (arxiv.org)

Mathematical objective

The original GAN defines a minimax game, connecting its training procedure to game theory:

min⁡Gmax⁡DV(D,G)=Ex∼pdata[log⁡D(x)]+Ez∼pz[log⁡(1−D(G(z)))].\min_G\max_D V(D,G) = \mathbb{E}_{x\sim p_{\mathrm{data}}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))].

Here, pdatap_{\mathrm{data}} is the real-data distribution, pzp_z is the input-noise distribution, and E\mathbb{E} denotes an expected value. The discriminator maximizes this objective, while the generator minimizes it. With unrestricted model capacity and an optimal discriminator, the theoretical optimum occurs when the generated and real distributions coincide; the discriminator then outputs 1/21/2 on their support. This idealized result does not guarantee convergence for finite neural networks. (arxiv.org)

In practice, the original generator loss function can provide weak gradients when the discriminator readily rejects generated samples. A frequently used alternative is the non-saturating loss, which minimizes −Ez[log⁡D(G(z))]-\mathbb{E}_z[\log D(G(z))]. It changes the generator’s learning signal while preserving the desired distribution-matching solution. Other GAN variants replace the classification objective with different measures of discrepancy between distributions. (arxiv.org)

Major variants and architectural developments

A conditional GAN, introduced in 2014, supplies additional information to both networks. For example, a class label can specify which digit the generator should produce. Conditioning distinguishes controlled generation from merely sampling the overall training distribution. (arxiv.org)

Deep convolutional GANs, or DCGANs, use convolutional neural networks and architectural constraints to support image generation and representation learning. Their significance includes demonstrating that adversarial training can learn useful visual features without category annotations. (arxiv.org)

The Wasserstein GAN replaces the probability-valued discriminator with a real-valued critic and uses an objective related to the Wasserstein distance between distributions. Its formulation requires a critic constrained by Lipschitz continuity. The gradient-penalty variant, WGAN-GP, introduced a penalty on the critic’s input-gradient norm as an alternative to clipping network weights. These changes address training difficulties but do not eliminate every failure case. (proceedings.mlr.press)

StyleGAN, first presented in December 2018, introduced a mapping network and layer-specific style controls. Its architecture separates aspects of large-scale image structure from stochastic details, enabling control at different synthesis scales. The original paper demonstrated this approach using generated human faces. (arxiv.org)

Training difficulties

GAN training is a coupled optimization problem: updating either network changes the objective faced by the other. Consequently, losses may oscillate, and improved discrimination need not correspond directly to improved generation. Network capacity, update schedules, and learning rates affect this interaction. Two-time-scale training assigns different learning rates to the networks; convergence results apply under specified mathematical assumptions rather than universally. (arxiv.org)

A characteristic failure is mode collapse, in which many inputs produce similar outputs, leaving substantial portions of the target distribution unrepresented. A generator may therefore produce convincing individual images while failing to capture the dataset’s diversity. Remedies studied include modified objectives, feature matching, minibatch-based discrimination, and regularization. No single technique guarantees both stable training and complete distribution coverage. (arxiv.org)

Applications and evaluation

Within computer vision, conditional GANs support image-to-image translation, including generating photographs from semantic label maps, reconstructing images from edges, and colorization. Higher-resolution systems also permit interactive changes to object categories and appearance through structured input controls. These tasks combine adversarial realism objectives with constraints linking outputs to their inputs. (arxiv.org)

Evaluation must distinguish sample quality from distribution coverage. Fréchet Inception Distance (FID) compares real and generated images through the means and covariances of features extracted by an Inception network; lower values indicate closer agreement under this representation. Precision-and-recall approaches separately assess fidelity and coverage, revealing differences that a single scalar score can conceal. Neither attractive examples nor one aggregate metric alone establishes that a generator faithfully represents all aspects of its training distribution. (arxiv.org)