# GANs Explained: How AI Learns to Create New Data

> Source: <https://pub.towardsai.net/gans-explained-how-ai-learns-to-create-new-data-9a8a124f77bd?source=rss----98111c9905da---4>
> Published: 2026-10-03 13:31:01+00:00

A **GAN** is two neural networks locked in a game. The **Generator** turns random noise into fake data; the **Discriminator** tries to tell fake from real. Each one’s mistakes teach the other, and over time the fakes become convincing. By the end of this post you’ll understand every part of that loop, the loss function, why GANs are hard to train, and you’ll have a working MNIST GAN in Keras.

Most neural networks people meet first have one job: **look at data and give an answer.**

A CNN looks at an image and says “cat.” An RNN or LSTM reads a sentence and says “positive.” Input in, label or prediction out.

Now for a twist. What if we don’t want the network to *answer* a question, but to *make* something? A face that belongs to no one. A handwritten digit nobody wrote. A painting no one painted.

That raises the question this whole post is about:

**How can a neural network create something that didn’t exist in the training data?**

The answer is one of the most creative ideas in deep learning: **make two networks compete.**

**What you’ll learn in this post:**

1. Predicting vs classifying vs generating

2. What a GAN is, and the counterfeiter-vs-detective analogy

3. The Generator and Discriminator, layer by layer

4. How random noise becomes an image

5. The GAN loss and the training loop, step by step

6. How the Generator actually learns (through the Discriminator)

7. GAN types, GAN vs CNN, and what goes wrong (mode collapse, instability)

8. A complete MNIST GAN in Keras

Let’s separate three things that often get blurred together:

Generative models have a different goal. They don’t just learn a boundary between categories; they learn *what the data looks like* well enough to produce new examples of it.

Some things that can be done this way:

• **Generating images**

• **Creating synthetic data** (useful when real data is scarce)

• **Generating faces**

• **Creating artwork**

The idea we’ll explore is the **Generative Adversarial Network (GAN)**.

**A GAN is a system containing two neural networks that compete with each other: a Generator and a Discriminator.**

My notes describe it as two networks *“trained simultaneously through adversarial training.”* The core flow is tiny:

Generator

 ↓

 Fake Data

 ↓

 Discriminator

 ↓

 Real or Fake?

Here it is with the real data added, plus the feedback loop that makes learning possible:

A GAN in one picture

Read it left to right: random noise goes into the Generator, which makes fake data. The Discriminator receives **both** fake data and real data and has to say which is which. The dashed arrow is the crucial part: the Discriminator’s verdict flows back as **feedback** that teaches the Generator.

• Takes **random noise** as input

• Produces data (like an image) **as close as possible to the real data**

• Its goal: make fakes so good that the Discriminator believes they’re real

• Takes **real data and generated data** as input

• Tries to tell them apart

• Outputs the **probability that a given sample is real**

Generator → Fake Data ──┐

 ↓

 Discriminator

 ↑

 Real Data ──────────────┘

Notice the Discriminator is just a **binary classifier**: “real” or “fake.” You already know how those work. The new thing is that its opponent is another network.

Imagine a counterfeiter printing fake money and a detective whose job is to spot it.

• **Generator = counterfeiter.** Prints fakes and tries to make them look genuine.

Counterfeiter vs detective

At the start, the counterfeiter is terrible and the detective catches everything. But every time the detective spots a fake, the counterfeiter learns *what gave it away* and improves. As the fakes get better, the detective has to get sharper too.

Generator gets better

 ↕

 Discriminator gets better

My notes put it this way: *the generator produces, the discriminator distinguishes, and this competing game generates better data over time.*

**Where the analogy breaks down:** a real counterfeiter might never see the detective’s reasoning. In a GAN, the Generator receives *mathematical feedback* (gradients) from the Discriminator, telling it precisely which direction to change its output. We’ll see this in section 12.

Here is one complete cycle of the loop. Most pieces will look familiar from how any neural network trains:

One GAN training cycle

Random Noise → Generator → Fake Data → Discriminator

 → Real or Fake? → Calculate Loss → Backpropagation

 → Update Networks → Repeat

It’s the same machinery as any neural network: a **forward pass**, a **loss**, **backpropagation**, and **gradient descent**. The twist is that there are **two networks**, each with its own loss, each trying to beat the other.

The Generator doesn’t start from a picture. It starts from a **random vector**, called a **latent vector** or **noise vector**:

[0.21, -0.73, 0.45, 0.91, …]

 ↓

 Generator

 ↓

 Generated Image

In my notes, this input is *sampled from a uniform distribution* and “forms the starting point of generating.” (Many implementations use a normal distribution instead. Either way, it’s just random numbers.)

**Why random?** Because different random inputs lead to **different outputs**. The noise is the source of variety. If you gave the Generator the same input every time, it would produce the same image every time. The Generator’s real job is to learn a transformation that turns meaningless random numbers into **structured data**.

Over training, the Generator learns the patterns that make real data look real. For images, it builds up from simple to complex, much like the CNN story told in reverse:

Pixels

 ↓

 Edges

 ↓

 Shapes

 ↓

 Textures

 ↓

 Objects

 ↓

 Realistic Image

Structurally, a typical image Generator looks like this (straight from my notes): an **input** noise vector, **fully connected layers** that reshape the noise, **batch normalization**, **activation functions** (ReLU / LeakyReLU), **transposed convolutional layers**, then **reshaping and an output layer**.

Inside the Generator

The diagram shows a small MNIST-sized Generator:

1. **100 random numbers** go in.

2. A **Dense** layer expands them, followed by **BatchNorm** and **LeakyReLU**.

3. **Reshape** turns the long vector into a small 7×7 feature map with 128 channels.

4. **Transposed convolutions** grow that into 14×14, then 28×28: the size of an MNIST digit.

5. The last layer uses **tanh**, so pixel values land between −1 and 1.

If a convolution *shrinks* an image into features (the way CNNs do), a **transposed convolution** does the opposite: it *grows* features back into an image.

The Discriminator learns to separate two groups:

Real Data ──────→ Real

 \

 → Discriminator

 /

 Generated Data ──→ Fake

It’s a CNN-style classifier. Per my notes, its typical structure is: **convolutional layers**, **batch normalization**, **LeakyReLU**, **pooling layers** (many modern designs use strided convolutions for downsampling instead), a **fully connected layer**, and an **output layer** giving a single value between 0 and 1.

Inside the Discriminator

**Why LeakyReLU instead of ReLU?** My notes explain it from the CNN/activation chapter: LeakyReLU fixes the “dead neuron” problem (ReLU outputs zero for every negative input), keeps a small gradient for negative values, and gives *faster and more stable training*. That matters a lot in GANs, where stable gradients are hard to come by.

This is the heart of the idea:

Generator:

 “Can I fool you?”

 

 VS

 

 Discriminator:

 “Can I detect you?”

During training:

• The Generator improves at creating realistic data

• The Discriminator improves at detecting generated data

• Both networks keep adapting to each other

Neither one is handed the “right answer” for what a good fake looks like. The Generator only learns because the Discriminator keeps pushing back.

Here’s what’s unusual: the two networks are optimized for **opposite objectives.**

**The Generator wants** the Discriminator to classify its generated data as *real*.

**The Discriminator wants** to correctly classify *both* real and generated data.

Two players, opposite goals

Both use **binary cross-entropy** to calculate their loss (my notes: “both G and D use binary cross entropy to calculate loss”). It’s the same loss we use for any yes/no classifier.

The original paper combines both goals into one “minimax” objective:

One score, two players pulling in opposite directions. (In practice, implementations often tweak the Generator’s loss slightly so it learns better early on. The idea is the same: the Generator is rewarded when the Discriminator is fooled.)

**Step 1:** Give random noise to the Generator

**Step 2:** The Generator creates fake data

**Step 3:** Give real and fake data to the Discriminator

**Step 4:** The Discriminator predicts real or fake

**Step 5:** Calculate both losses

**Step 6:** Update the Discriminator

**Step 7:** Update the Generator

**Step 8:** Repeat

In steps 6 and 7 the two networks take turns. While one is being updated, the other is held still. (In the Keras code later, we’ll compute both gradients from the same forward pass and apply them together, which is a common simplification.)

The Generator never sees a real image as its target. So how does it learn?

Bad Image

 ↓

 Discriminator detects fake

 ↓

 Generator receives gradient

 ↓

 Generator updates weights

 ↓

 Better Image

 ↓

 Repeat

**The Generator improves because the Discriminator provides a learning signal.**

Gradient flow

During the **forward pass**, noise travels through the Generator, becomes a fake image, passes through the Discriminator, and produces a loss. During the **backward pass**, the gradient flows from the loss, *through the Discriminator*, and into the Generator. It tells the Generator how to change each pixel so the Discriminator’s score goes up.

That’s ordinary backpropagation and gradient descent, now passing through two networks chained together. The Discriminator is updated to catch fakes better; the Generator is updated to fool it.

Once training is done, you don’t need the Discriminator anymore. Generating is just:

Random Vector

 ↓

 Generator

 ↓

 Learned Features

 ↓

 Image

Random Noise

 ↓

 GAN

 ↓

 New Face

Each new random vector gives a new output. The generated face is a **new synthetic example**, not a copy-paste of one training photo. It’s a fresh sample drawn from what the Generator learned about the data.

My notes list these applications:

1. **Image generation**

2. **Data augmentation** (more training examples when real data is limited)

3. **Style transfer**

4. **Super-resolution** (sharpening or enlarging images)

5. **Generating art**

6. **Image-to-image translation**

Other uses include **synthetic datasets** and research in **synthetic media**.

A note on responsible use: the same technology that makes artwork and helpful training data can also create **misleading synthetic media**. Labeling generated content and using it honestly matters.

There’s a whole family. Here’s the quick tour, based on my notes:

The GAN family

**Vanilla GAN.** The simplest form, built using MLP layers. It’s the foundation of everything else. Known for training instability and mode collapse.

**DCGAN (Deep Convolutional GAN).** Replaces fully connected layers with convolutional ones. It’s highly stable, produces high-quality, sharp images, and is a common choice for image generation.

**Conditional GAN (cGAN).** Adds conditional variables like labels, attributes or text, so you can *command* the generator to produce a certain output.

Label = “Dog”

 +

 Random Noise

 ↓

 Generator

 ↓

 Dog Image

**CycleGAN.** Image-to-image translation that **doesn’t need matching pairs** of training examples.

Horse → Zebra

**StyleGAN.** Produces photorealistic, high-resolution faces with hierarchical control over “style.” Developed by NVIDIA.

16. GAN vs a traditional neural network

People sometimes ask “which is better, a GAN or a CNN?” That’s a bit like asking whether a *recipe* is better than an *oven*.

• **CNN** is an **architecture** commonly used to process images.

GAN vs CNN

For image tasks, a GAN’s Generator and Discriminator can themselves be **CNN-based**. That’s exactly what DCGAN does.

GANs are impressive and famously finicky. Here are the big three.

The Generator finds a few outputs that reliably fool the Discriminator and produces *only* those, instead of the full variety of the data.

Mode collapse

Expected: Generated:

 Cat A Cat A

 Cat B Cat A

 Cat C Cat A

 Cat D Cat A

One mitigation from my notes: **minibatch** techniques, which let the Discriminator look at groups of samples rather than one at a time, so it can notice when everything looks the same.

Training two networks at once is hard: they’re two models of a neural network, and the optimization landscape is tricky. They can become unbalanced:

Training instability

• **An overpowering Discriminator.** It gets so good that its strong output gives the Generator almost nothing useful to learn from. Helpful tricks: **label smoothing**, **feature matching**, or **slowing down** the Discriminator’s training.

• **A weak Discriminator.** It can’t give the Generator meaningful feedback to improve.

When things go wrong, the usual advice is to choose a GAN architecture that suits the problem and rethink the optimization setup.

How do you score a “good” face or digit? Measuring both the **quality** and the **diversity** of generated data isn’t straightforward, and often comes down to a mix of metrics and human judgment.

MNIST (handwritten digits) is the ideal beginner dataset, because you can *see* whether the output looks like a digit.

Load Dataset → Build Generator → Build Discriminator → Generate Fake Images

 → Train Discriminator → Train Generator → Repeat

**Setup and data**

import tensorflow as tf

 from tensorflow.keras import layers, models

 import matplotlib.pyplot as plt

 

 NOISE_DIM = 100

 BATCH_SIZE = 128

 

 (x_train, _), _ = tf.keras.datasets.mnist.load_data()

 x_train = x_train.astype(“float32”)

 x_train = (x_train — 127.5) / 127.5 # scale pixels to [-1, 1]

 x_train = x_train[…, None] # shape: (60000, 28, 28, 1)

 

 dataset = (

 tf.data.Dataset.from_tensor_slices(x_train)

 .shuffle(60000)

 .batch(BATCH_SIZE, drop_remainder=True)

 )

**The Generator:** noise in, 28×28 image out.

def build_generator():

 return models.Sequential([

 layers.Input(shape=(NOISE_DIM,)),

 layers.Dense(7 * 7 * 128, use_bias=False),

 layers.BatchNormalization(),

 layers.LeakyReLU(0.2),

 layers.Reshape((7, 7, 128)),

 layers.Conv2DTranspose(64, 4, strides=2, padding=”same”, use_bias=False), # 14x14

 layers.BatchNormalization(),

 layers.LeakyReLU(0.2),

 layers.Conv2DTranspose(1, 4, strides=2, padding=”same”, activation=”tanh”), # 28x28

 ])

**The Discriminator:** image in, one “realness” score out.

def build_discriminator():

 return models.Sequential([

 layers.Input(shape=(28, 28, 1)),

 layers.Conv2D(64, 4, strides=2, padding=”same”),

 layers.LeakyReLU(0.2),

 layers.Dropout(0.3),

 layers.Conv2D(128, 4, strides=2, padding=”same”),

 layers.BatchNormalization(),

 layers.LeakyReLU(0.2),

 layers.Dropout(0.3),

 layers.Flatten(),

 layers.Dense(1), # raw score (logit)

 ])

 

 generator = build_generator()

 discriminator = build_discriminator()

**Losses and optimizers**

bce = tf.keras.losses.BinaryCrossentropy(from_logits=True)

 

 def discriminator_loss(real_out, fake_out):

 real_loss = bce(tf.ones_like(real_out), real_out) # real should score 1

 fake_loss = bce(tf.zeros_like(fake_out), fake_out) # fake should score 0

 return real_loss + fake_loss

 

 def generator_loss(fake_out):

 return bce(tf.ones_like(fake_out), fake_out) # G wants fakes scored as real

 

 g_opt = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)

 d_opt = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)

**One training step:** the 8-step cycle from earlier, in code.

@tf.function

 def train_step(real_images):

 noise = tf.random.normal([BATCH_SIZE, NOISE_DIM])

 

 with tf.GradientTape() as g_tape, tf.GradientTape() as d_tape:

 fake_images = generator(noise, training=True) # steps 1–2

 

 real_out = discriminator(real_images, training=True) # steps 3–4

 fake_out = discriminator(fake_images, training=True)

 

 g_loss = generator_loss(fake_out) # step 5

 d_loss = discriminator_loss(real_out, fake_out)

 

 d_grads = d_tape.gradient(d_loss, discriminator.trainable_variables)

 g_grads = g_tape.gradient(g_loss, generator.trainable_variables)

 d_opt.apply_gradients(zip(d_grads, discriminator.trainable_variables)) # step 6

 g_opt.apply_gradients(zip(g_grads, generator.trainable_variables)) # step 7

 return g_loss, d_loss

**Training loop that saves samples so you can watch it learn**

seed = tf.random.normal([16, NOISE_DIM]) # same 16 noise vectors every time

 

 def save_samples(epoch):

 images = generator(seed, training=False)

 fig, axes = plt.subplots(4, 4, figsize=(4, 4))

 for ax, img in zip(axes.flat, images):

 ax.imshow(img[:, :, 0] * 0.5 + 0.5, cmap=”gray”) # back to [0, 1]

 ax.axis(“off”)

 plt.savefig(f”epoch_{epoch:03d}.png”, bbox_inches=”tight”)

 plt.close(fig)

 

 EPOCHS = 50

 for epoch in range(1, EPOCHS + 1):

 for batch in dataset:

 g_loss, d_loss = train_step(batch)

 print(f”Epoch {epoch}: G loss {g_loss:.3f} | D loss {d_loss:.3f}”)

 if epoch in (1, 10, 25, 50):

 save_samples(epoch)

Because we reuse the **same** seed every time, the 16 images in each saved grid are directly comparable: it’s the same 16 noise vectors being interpreted by an increasingly skilled Generator.

This is the most satisfying part of building a GAN: **watching the Generator improve.**

Epoch 1 → Random-looking images

 Epoch 10 → Basic shapes

 Epoch 25 → Recognizable patterns

 Epoch 50 → More realistic samples

Epoch template

📌 **Your turn:** run the code above, then place your four saved grids (epoch_001.png, epoch_010.png, epoch_025.png, epoch_050.png) side by side where the template shows. The exact look depends on your run and settings, so use your own real outputs rather than the rough pattern above.

Early on you’ll see static-like blobs. Later, stroke-like shapes appear, and eventually many of the samples start to look like handwritten digits.

1. **Thinking the Generator sees real images.** It never does; it only receives gradients through the Discriminator.

2. **Letting the Discriminator win too easily.** If its loss drops to ~0 quickly, the Generator may stop learning.

3. **Watching only the loss.** In GANs, a falling loss doesn’t guarantee better images. Look at the samples.

4. **Forgetting to scale pixels** to match the Generator’s tanh output range (−1 to 1).

5. **Expecting stable training.** Run-to-run variation is normal; tuning learning rates and architecture matters.

6. **Treating “GAN” and “CNN” as alternatives.** One is a framework; the other an architecture.

**What does GAN stand for?** Generative Adversarial Network.

**Why does a GAN need two networks?** One network creates and the other judges. The judge’s feedback is what teaches the creator. Without an opponent, the Generator has no signal about what “realistic” means.

**What is the Generator?** The network that turns random noise into synthetic data.

**What is the Discriminator?** A binary classifier that decides whether a sample is real or generated.

**What is random noise?** A random vector (the latent vector) that gives the Generator a starting point and is the source of variety in the outputs.

**How does the Generator learn?** From gradients that flow back through the Discriminator. They show how to change the output to look more real.

**How does the Discriminator learn?** Like any classifier: it’s shown real samples labeled “real” and generated samples labeled “fake,” and it minimizes its classification loss.

**What is mode collapse?** When the Generator produces only a narrow range of outputs instead of the full variety in the data.

**Can GANs create completely new data?** They create new samples that follow the patterns they’ve learned, and generally aren’t copies of individual training examples. But they’re limited by what they learned from: “new” here means new combinations within the learned patterns, not invention from nothing.

**GAN vs CNN?** A CNN is an architecture. A GAN is a framework that can use CNNs inside both of its networks.

**GAN vs traditional neural network?** Traditional networks map input to a prediction. A GAN pairs two networks to learn to generate realistic samples.

**Aren’t autoencoders generative too?** A plain autoencoder compresses data into a latent space and reconstructs it, which is great for dimensionality reduction, denoising and anomaly detection, but it’s not built to generate new content. **VAEs (Variational Autoencoders)** go a step further: the encoder outputs the parameters of a probability distribution, we sample from it, and the decoder turns the sample into data. That makes them generative. It’s another road to the same destination, which we won’t go down today.

Random Noise

 ↓

 Generator

 ↓

 Fake Data

 ↓

 Discriminator ← Real Data

 ↓

 Feedback

 ↓

 Improve Generator

 ↓

 More Realistic Data

• A GAN is **two networks**: a Generator (the forger) and a Discriminator (the detective).

• The Generator turns **random noise** into data; the Discriminator is a **binary classifier**.

• They have **opposite goals** and both use **binary cross-entropy**.

• The Generator learns from **gradients that flow back through the Discriminator**.

• Watch out for **mode collapse** and **training instability**.

• GANs are a *framework*; they often use **CNNs** inside.

**A GAN learns through competition: the Generator learns to create increasingly realistic data, while the Discriminator learns to distinguish generated data from real data.**

GANs showed how neural networks can **generate** data by competing. Good next steps: train the MNIST GAN above and save your own epoch-by-epoch samples, try a **conditional GAN** so you can ask for a specific digit, and compare GANs with other generative approaches such as **VAEs (Variational Autoencoders)** and Transformer-based generators.

*If this post helped, a clap or a follow means a lot, and I’d love to see your own epoch-by-epoch samples in the comments.*

[GANs Explained: How AI Learns to Create New Data](https://pub.towardsai.net/gans-explained-how-ai-learns-to-create-new-data-9a8a124f77bd) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
