cd /news/artificial-intelligence/gans-explained-how-ai-learns-to-crea… · home › topics › artificial-intelligence › article
[ARTICLE · art-144475] src=pub.towardsai.net ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

GANs Explained: How AI Learns to Create New Data

A Generative Adversarial Network (GAN) consists of two neural networks — a Generator that turns random noise into fake data and a Discriminator that outputs the probability a sample is real — trained simultaneously through adversarial training so each network's mistakes teach the other. The explainer walks through the GAN loss and training loop step by step, covers failure modes including mode collapse and instability, and includes a complete MNIST GAN implemented in Keras.

by read17 min views1 publishedOct 3, 2026

A GAN is two neural networks locked in a game. The Generator turns random noise into fake data; the Discriminator tries to tell fake from real. Each one’s mistakes teach the other, and over time the fakes become convincing. By the end of this post you’ll understand every part of that loop, the loss function, why GANs are hard to train, and you’ll have a working MNIST GAN in Keras.

Most neural networks people meet first have one job: look at data and give an answer.

A CNN looks at an image and says “cat.” An RNN or LSTM reads a sentence and says “positive.” Input in, label or prediction out.

Now for a twist. What if we don’t want the network to answer a question, but to make something? A face that belongs to no one. A handwritten digit nobody wrote. A painting no one painted.

That raises the question this whole post is about:

How can a neural network create something that didn’t exist in the training data?

The answer is one of the most creative ideas in deep learning: make two networks compete.

What you’ll learn in this post:

  1. Predicting vs classifying vs generating

  2. What a GAN is, and the counterfeiter-vs-detective analogy

  3. The Generator and Discriminator, layer by layer

  4. How random noise becomes an image

  5. The GAN loss and the training loop, step by step

  6. How the Generator actually learns (through the Discriminator)

  7. GAN types, GAN vs CNN, and what goes wrong (mode collapse, instability)

  8. A complete MNIST GAN in Keras

Let’s separate three things that often get blurred together:

Generative models have a different goal. They don’t just learn a boundary between categories; they learn what the data looks like well enough to produce new examples of it.

Some things that can be done this way:

• Generating images

• Creating synthetic data (useful when real data is scarce)

• Generating faces

• Creating artwork

The idea we’ll explore is the Generative Adversarial Network (GAN).

A GAN is a system containing two neural networks that compete with each other: a Generator and a Discriminator.

My notes describe it as two networks “trained simultaneously through adversarial training.” The core flow is tiny:

Generator

↓

Fake Data

↓

Discriminator

↓

Real or Fake?

Here it is with the real data added, plus the feedback loop that makes learning possible:

A GAN in one picture

Read it left to right: random noise goes into the Generator, which makes fake data. The Discriminator receives both fake data and real data and has to say which is which. The dashed arrow is the crucial part: the Discriminator’s verdict flows back as feedback that teaches the Generator.

• Takes random noise as input

• Produces data (like an image) as close as possible to the real data

• Its goal: make fakes so good that the Discriminator believes they’re real

• Takes real data and generated data as input

• Tries to tell them apart

• Outputs the probability that a given sample is real

Generator → Fake Data ──┐

↓

Discriminator

↑

Real Data ──────────────┘

Notice the Discriminator is just a binary classifier: “real” or “fake.” You already know how those work. The new thing is that its opponent is another network.

Imagine a counterfeiter printing fake money and a detective whose job is to spot it.

• Generator = counterfeiter. Prints fakes and tries to make them look genuine.

Counterfeiter vs detective

At the start, the counterfeiter is terrible and the detective catches everything. But every time the detective spots a fake, the counterfeiter learns what gave it away and improves. As the fakes get better, the detective has to get sharper too.

Generator gets better

↕

Discriminator gets better

My notes put it this way: the generator produces, the discriminator distinguishes, and this competing game generates better data over time.

Where the analogy breaks down: a real counterfeiter might never see the detective’s reasoning. In a GAN, the Generator receives mathematical feedback (gradients) from the Discriminator, telling it precisely which direction to change its output. We’ll see this in section 12.

Here is one complete cycle of the loop. Most pieces will look familiar from how any neural network trains:

One GAN training cycle

Random Noise → Generator → Fake Data → Discriminator

→ Real or Fake? → Calculate Loss → Backpropagation

→ Update Networks → Repeat

It’s the same machinery as any neural network: a forward pass, a loss, backpropagation, and gradient descent. The twist is that there are two networks, each with its own loss, each trying to beat the other.

The Generator doesn’t start from a picture. It starts from a random vector, called a latent vector or noise vector:

[0.21, -0.73, 0.45, 0.91, …] ↓

Generator

↓

Generated Image

In my notes, this input is sampled from a uniform distribution and “forms the starting point of generating.” (Many implementations use a normal distribution instead. Either way, it’s just random numbers.)

Why random? Because different random inputs lead to different outputs. The noise is the source of variety. If you gave the Generator the same input every time, it would produce the same image every time. The Generator’s real job is to learn a transformation that turns meaningless random numbers into structured data.

Over training, the Generator learns the patterns that make real data look real. For images, it builds up from simple to complex, much like the CNN story told in reverse:

Pixels

↓

Edges

↓

Shapes

↓

Textures

↓

Objects

↓

Realistic Image

Structurally, a typical image Generator looks like this (straight from my notes): an input noise vector, fully connected layers that reshape the noise, batch normalization, activation functions (ReLU / LeakyReLU), transposed convolutional layers, then reshaping and an output layer.

Inside the Generator

The diagram shows a small MNIST-sized Generator:

  1. 100 random numbers go in.

  2. A Dense layer expands them, followed by BatchNorm and LeakyReLU.

  3. Reshape turns the long vector into a small 7×7 feature map with 128 channels.

  4. Transposed convolutions grow that into 14×14, then 28×28: the size of an MNIST digit.

  5. The last layer uses tanh, so pixel values land between −1 and 1.

If a convolution shrinks an image into features (the way CNNs do), a transposed convolution does the opposite: it grows features back into an image. The Discriminator learns to separate two groups:

Real Data ──────→ Real

\

→ Discriminator

/

Generated Data ──→ Fake

It’s a CNN-style classifier. Per my notes, its typical structure is: convolutional layers, batch normalization, LeakyReLU, pooling layers (many modern designs use strided convolutions for downsampling instead), a fully connected layer, and an output layer giving a single value between 0 and 1.

Inside the Discriminator

Why LeakyReLU instead of ReLU? My notes explain it from the CNN/activation chapter: LeakyReLU fixes the “dead neuron” problem (ReLU outputs zero for every negative input), keeps a small gradient for negative values, and gives faster and more stable training. That matters a lot in GANs, where stable gradients are hard to come by.

This is the heart of the idea:

Generator:

“Can I fool you?”

VS

Discriminator:

“Can I detect you?”

During training:

• The Generator improves at creating realistic data

• The Discriminator improves at detecting generated data

• Both networks keep adapting to each other

Neither one is handed the “right answer” for what a good fake looks like. The Generator only learns because the Discriminator keeps pushing back.

Here’s what’s unusual: the two networks are optimized for opposite objectives.

The Generator wants the Discriminator to classify its generated data as real.

The Discriminator wants to correctly classify both real and generated data.

Two players, opposite goals

Both use binary cross-entropy to calculate their loss (my notes: “both G and D use binary cross entropy to calculate loss”). It’s the same loss we use for any yes/no classifier.

The original paper combines both goals into one “minimax” objective:

One score, two players pulling in opposite directions. (In practice, implementations often tweak the Generator’s loss slightly so it learns better early on. The idea is the same: the Generator is rewarded when the Discriminator is fooled.)

Step 1: Give random noise to the Generator

Step 2: The Generator creates fake data

Step 3: Give real and fake data to the Discriminator

Step 4: The Discriminator predicts real or fake

Step 5: Calculate both losses

Step 6: Update the Discriminator

Step 7: Update the Generator

Step 8: Repeat

In steps 6 and 7 the two networks take turns. While one is being updated, the other is held still. (In the Keras code later, we’ll compute both gradients from the same forward pass and apply them together, which is a common simplification.)

The Generator never sees a real image as its target. So how does it learn?

Bad Image

↓

Discriminator detects fake

↓

Generator receives gradient

↓

Generator updates weights

↓

Better Image

↓

Repeat

The Generator improves because the Discriminator provides a learning signal.

Gradient flow

During the forward pass, noise travels through the Generator, becomes a fake image, passes through the Discriminator, and produces a loss. During the backward pass, the gradient flows from the loss, through the Discriminator, and into the Generator. It tells the Generator how to change each pixel so the Discriminator’s score goes up.

That’s ordinary backpropagation and gradient descent, now passing through two networks chained together. The Discriminator is updated to catch fakes better; the Generator is updated to fool it.

Once training is done, you don’t need the Discriminator anymore. Generating is just:

Random Vector

↓

Generator

↓

Learned Features

↓

Image

Random Noise

↓

GAN

↓

New Face Each new random vector gives a new output. The generated face is a new synthetic example, not a copy-paste of one training photo. It’s a fresh sample drawn from what the Generator learned about the data.

My notes list these applications:

  1. Image generation

  2. Data augmentation (more training examples when real data is limited)

  3. Style transfer

  4. Super-resolution (sharpening or enlarging images)

  5. Generating art

  6. Image-to-image translation Other uses include synthetic datasets and research in synthetic media.

A note on responsible use: the same technology that makes artwork and helpful training data can also create misleading synthetic media. Labeling generated content and using it honestly matters.

There’s a whole family. Here’s the quick tour, based on my notes:

The GAN family

Vanilla GAN. The simplest form, built using MLP layers. It’s the foundation of everything else. Known for training instability and mode collapse.

DCGAN (Deep Convolutional GAN). Replaces fully connected layers with convolutional ones. It’s highly stable, produces high-quality, sharp images, and is a common choice for image generation.

Conditional GAN (cGAN). Adds conditional variables like labels, attributes or text, so you can command the generator to produce a certain output.

Label = “Dog”

Random Noise

↓

Generator

↓

Dog Image

CycleGAN. Image-to-image translation that doesn’t need matching pairs of training examples.

Horse → Zebra

StyleGAN. Produces photorealistic, high-resolution faces with hierarchical control over “style.” Developed by NVIDIA.

  1. GAN vs a traditional neural network

People sometimes ask “which is better, a GAN or a CNN?” That’s a bit like asking whether a recipe is better than an oven.

• CNN is an architecture commonly used to process images.

GAN vs CNN

For image tasks, a GAN’s Generator and Discriminator can themselves be CNN-based. That’s exactly what DCGAN does. GANs are impressive and famously finicky. Here are the big three.

The Generator finds a few outputs that reliably fool the Discriminator and produces only those, instead of the full variety of the data.

Mode collapse

Expected: Generated: Cat A Cat A

Cat B Cat A

Cat C Cat A

Cat D Cat A

One mitigation from my notes: minibatch techniques, which let the Discriminator look at groups of samples rather than one at a time, so it can notice when everything looks the same.

Training two networks at once is hard: they’re two models of a neural network, and the optimization landscape is tricky. They can become unbalanced:

Training instability

• An overpowering Discriminator. It gets so good that its strong output gives the Generator almost nothing useful to learn from. Helpful tricks: label smoothing, feature matching, or slowing down the Discriminator’s training.

• A weak Discriminator. It can’t give the Generator meaningful feedback to improve.

When things go wrong, the usual advice is to choose a GAN architecture that suits the problem and rethink the optimization setup.

How do you score a “good” face or digit? Measuring both the quality and the diversity of generated data isn’t straightforward, and often comes down to a mix of metrics and human judgment.

MNIST (handwritten digits) is the ideal beginner dataset, because you can see whether the output looks like a digit.

Load Dataset → Build Generator → Build Discriminator → Generate Fake Images

→ Train Discriminator → Train Generator → Repeat

Setup and data

import tensorflow as tf

 from tensorflow.keras import layers, models

 import matplotlib.pyplot as plt

NOISE_DIM = 100

BATCH_SIZE = 128

 (x_train, _), _ = tf.keras.datasets.mnist.load_data()

 x_train = x_train.astype(“float32”)

 x_train = (x_train — 127.5) / 127.5 # scale pixels to [-1, 1]

 x_train = x_train[…, None] # shape: (60000, 28, 28, 1)

 

 dataset = (

tf.data.Dataset.from_tensor_slices(x_train)

 .shuffle(60000)

 .batch(BATCH_SIZE, drop_remainder=True)

 )

The Generator: noise in, 28×28 image out.

def build_generator():

 return models.Sequential([

 layers.Input(shape=(NOISE_DIM,)),

 layers.Dense(7 * 7 * 128, use_bias=False),

 layers.BatchNormalization(),

 layers.LeakyReLU(0.2),

 layers.Reshape((7, 7, 128)),

 layers.Conv2DTranspose(64, 4, strides=2, padding=”same”, use_bias=False), # 14x14

 layers.BatchNormalization(),

 layers.LeakyReLU(0.2),

 layers.Conv2DTranspose(1, 4, strides=2, padding=”same”, activation=”tanh”), # 28x28

 ])

The Discriminator: image in, one “realness” score out.

def build_discriminator():

 return models.Sequential([

 layers.Input(shape=(28, 28, 1)),

 layers.Conv2D(64, 4, strides=2, padding=”same”),

 layers.LeakyReLU(0.2),

 layers.Dropout(0.3),

 layers.Conv2D(128, 4, strides=2, padding=”same”),

 layers.BatchNormalization(),

 layers.LeakyReLU(0.2),

 layers.Dropout(0.3),

 layers.Flatten(),

 layers.Dense(1), # raw score (logit)

 ])

 

 generator = build_generator()

 discriminator = build_discriminator()

Losses and optimizers

bce = tf.keras.losses.BinaryCrossentropy(from_logits=True)

 

 def discriminator_loss(real_out, fake_out):

 real_loss = bce(tf.ones_like(real_out), real_out) # real should score 1

 fake_loss = bce(tf.zeros_like(fake_out), fake_out) # fake should score 0

 return real_loss + fake_loss

 

 def generator_loss(fake_out):

 return bce(tf.ones_like(fake_out), fake_out) # G wants fakes scored as real

 

 g_opt = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)

 d_opt = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)

One training step: the 8-step cycle from earlier, in code.

@tf.function

 def train_step(real_images):

 noise = tf.random.normal([BATCH_SIZE, NOISE_DIM])

 

 with tf.GradientTape() as g_tape, tf.GradientTape() as d_tape:

 fake_images = generator(noise, training=True) # steps 1–2

 

 real_out = discriminator(real_images, training=True) # steps 3–4

 fake_out = discriminator(fake_images, training=True)

 

 g_loss = generator_loss(fake_out) # step 5

 d_loss = discriminator_loss(real_out, fake_out)

d_grads = d_tape.gradient(d_loss, discriminator.trainable_variables)

g_grads = g_tape.gradient(g_loss, generator.trainable_variables)

d_opt.apply_gradients(zip(d_grads, discriminator.trainable_variables)) # step 6

 g_opt.apply_gradients(zip(g_grads, generator.trainable_variables)) # step 7

 return g_loss, d_loss

Training loop that saves samples so you can watch it learn

seed = tf.random.normal([16, NOISE_DIM]) # same 16 noise vectors every time

 

 def save_samples(epoch):

 images = generator(seed, training=False)

 fig, axes = plt.subplots(4, 4, figsize=(4, 4))

 for ax, img in zip(axes.flat, images):

 ax.imshow(img[:, :, 0] * 0.5 + 0.5, cmap=”gray”) # back to [0, 1]

 ax.axis(“off”)

 plt.savefig(f”epoch_{epoch:03d}.png”, bbox_inches=”tight”)

 plt.close(fig)

EPOCHS = 50

 for epoch in range(1, EPOCHS + 1):

 for batch in dataset:

 g_loss, d_loss = train_step(batch)

 print(f”Epoch {epoch}: G loss {g_loss:.3f} | D loss {d_loss:.3f}”)

 if epoch in (1, 10, 25, 50):

 save_samples(epoch)

Because we reuse the same seed every time, the 16 images in each saved grid are directly comparable: it’s the same 16 noise vectors being interpreted by an increasingly skilled Generator.

This is the most satisfying part of building a GAN: watching the Generator improve.

Epoch 1 → Random-looking images

Epoch 10 → Basic shapes

Epoch 25 → Recognizable patterns

Epoch 50 → More realistic samples

Epoch template

📌 Your turn: run the code above, then place your four saved grids (epoch_001.png, epoch_010.png, epoch_025.png, epoch_050.png) side by side where the template shows. The exact look depends on your run and settings, so use your own real outputs rather than the rough pattern above.

Early on you’ll see static-like blobs. Later, stroke-like shapes appear, and eventually many of the samples start to look like handwritten digits.

  1. Thinking the Generator sees real images. It never does; it only receives gradients through the Discriminator.

  2. Letting the Discriminator win too easily. If its loss drops to ~0 quickly, the Generator may stop learning.

  3. Watching only the loss. In GANs, a falling loss doesn’t guarantee better images. Look at the samples.

  4. Forgetting to scale pixels to match the Generator’s tanh output range (−1 to 1).

  5. Expecting stable training. Run-to-run variation is normal; tuning learning rates and architecture matters.

  6. Treating “GAN” and “CNN” as alternatives. One is a framework; the other an architecture.

What does GAN stand for? Generative Adversarial Network.

Why does a GAN need two networks? One network creates and the other judges. The judge’s feedback is what teaches the creator. Without an opponent, the Generator has no signal about what “realistic” means.

What is the Generator? The network that turns random noise into synthetic data.

What is the Discriminator? A binary classifier that decides whether a sample is real or generated.

What is random noise? A random vector (the latent vector) that gives the Generator a starting point and is the source of variety in the outputs.

How does the Generator learn? From gradients that flow back through the Discriminator. They show how to change the output to look more real.

How does the Discriminator learn? Like any classifier: it’s shown real samples labeled “real” and generated samples labeled “fake,” and it minimizes its classification loss.

What is mode collapse? When the Generator produces only a narrow range of outputs instead of the full variety in the data.

Can GANs create completely new data? They create new samples that follow the patterns they’ve learned, and generally aren’t copies of individual training examples. But they’re limited by what they learned from: “new” here means new combinations within the learned patterns, not invention from nothing.

GAN vs CNN? A CNN is an architecture. A GAN is a framework that can use CNNs inside both of its networks.

GAN vs traditional neural network? Traditional networks map input to a prediction. A GAN pairs two networks to learn to generate realistic samples.

Aren’t autoencoders generative too? A plain autoencoder compresses data into a latent space and reconstructs it, which is great for dimensionality reduction, denoising and anomaly detection, but it’s not built to generate new content. VAEs (Variational Autoencoders) go a step further: the encoder outputs the parameters of a probability distribution, we sample from it, and the decoder turns the sample into data. That makes them generative. It’s another road to the same destination, which we won’t go down today.

Random Noise

↓

Generator

↓

Fake Data

↓

Discriminator ← Real Data

↓

Feedback

↓

Improve Generator

↓

More Realistic Data

• A GAN is two networks: a Generator (the forger) and a Discriminator (the detective). • The Generator turns random noise into data; the Discriminator is a binary classifier.

• They have opposite goals and both use binary cross-entropy.

• The Generator learns from gradients that flow back through the Discriminator.

• Watch out for mode collapse and training instability.

• GANs are a framework; they often use CNNs inside.

A GAN learns through competition: the Generator learns to create increasingly realistic data, while the Discriminator learns to distinguish generated data from real data.

GANs showed how neural networks can generate data by competing. Good next steps: train the MNIST GAN above and save your own epoch-by-epoch samples, try a conditional GAN so you can ask for a specific digit, and compare GANs with other generative approaches such as VAEs (Variational Autoencoders) and Transformer-based generators.

If this post helped, a clap or a follow means a lot, and I’d love to see your own epoch-by-epoch samples in the comments.

GANs Explained: How AI Learns to Create New Data was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @generative adversarial network 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gans-explained-how-a…] indexed:0 read:17min 2026-10-03 · —