{"slug": "gans-explained-how-ai-learns-to-create-new-data", "title": "GANs Explained: How AI Learns to Create New Data", "summary": "A Generative Adversarial Network (GAN) consists of two neural networks — a Generator that turns random noise into fake data and a Discriminator that outputs the probability a sample is real — trained simultaneously through adversarial training so each network's mistakes teach the other. The explainer walks through the GAN loss and training loop step by step, covers failure modes including mode collapse and instability, and includes a complete MNIST GAN implemented in Keras.", "body_md": "A **GAN** is two neural networks locked in a game. The **Generator** turns random noise into fake data; the **Discriminator** tries to tell fake from real. Each one’s mistakes teach the other, and over time the fakes become convincing. By the end of this post you’ll understand every part of that loop, the loss function, why GANs are hard to train, and you’ll have a working MNIST GAN in Keras.\n\nMost neural networks people meet first have one job: **look at data and give an answer.**\n\nA CNN looks at an image and says “cat.” An RNN or LSTM reads a sentence and says “positive.” Input in, label or prediction out.\n\nNow for a twist. What if we don’t want the network to *answer* a question, but to *make* something? A face that belongs to no one. A handwritten digit nobody wrote. A painting no one painted.\n\nThat raises the question this whole post is about:\n\n**How can a neural network create something that didn’t exist in the training data?**\n\nThe answer is one of the most creative ideas in deep learning: **make two networks compete.**\n\n**What you’ll learn in this post:**\n\n1. Predicting vs classifying vs generating\n\n2. What a GAN is, and the counterfeiter-vs-detective analogy\n\n3. The Generator and Discriminator, layer by layer\n\n4. How random noise becomes an image\n\n5. The GAN loss and the training loop, step by step\n\n6. How the Generator actually learns (through the Discriminator)\n\n7. GAN types, GAN vs CNN, and what goes wrong (mode collapse, instability)\n\n8. A complete MNIST GAN in Keras\n\nLet’s separate three things that often get blurred together:\n\nGenerative models have a different goal. They don’t just learn a boundary between categories; they learn *what the data looks like* well enough to produce new examples of it.\n\nSome things that can be done this way:\n\n• **Generating images**\n\n• **Creating synthetic data** (useful when real data is scarce)\n\n• **Generating faces**\n\n• **Creating artwork**\n\nThe idea we’ll explore is the **Generative Adversarial Network (GAN)**.\n\n**A GAN is a system containing two neural networks that compete with each other: a Generator and a Discriminator.**\n\nMy notes describe it as two networks *“trained simultaneously through adversarial training.”* The core flow is tiny:\n\nGenerator\n\n ↓\n\n Fake Data\n\n ↓\n\n Discriminator\n\n ↓\n\n Real or Fake?\n\nHere it is with the real data added, plus the feedback loop that makes learning possible:\n\nA GAN in one picture\n\nRead it left to right: random noise goes into the Generator, which makes fake data. The Discriminator receives **both** fake data and real data and has to say which is which. The dashed arrow is the crucial part: the Discriminator’s verdict flows back as **feedback** that teaches the Generator.\n\n• Takes **random noise** as input\n\n• Produces data (like an image) **as close as possible to the real data**\n\n• Its goal: make fakes so good that the Discriminator believes they’re real\n\n• Takes **real data and generated data** as input\n\n• Tries to tell them apart\n\n• Outputs the **probability that a given sample is real**\n\nGenerator → Fake Data ──┐\n\n ↓\n\n Discriminator\n\n ↑\n\n Real Data ──────────────┘\n\nNotice the Discriminator is just a **binary classifier**: “real” or “fake.” You already know how those work. The new thing is that its opponent is another network.\n\nImagine a counterfeiter printing fake money and a detective whose job is to spot it.\n\n• **Generator = counterfeiter.** Prints fakes and tries to make them look genuine.\n\nCounterfeiter vs detective\n\nAt the start, the counterfeiter is terrible and the detective catches everything. But every time the detective spots a fake, the counterfeiter learns *what gave it away* and improves. As the fakes get better, the detective has to get sharper too.\n\nGenerator gets better\n\n ↕\n\n Discriminator gets better\n\nMy notes put it this way: *the generator produces, the discriminator distinguishes, and this competing game generates better data over time.*\n\n**Where the analogy breaks down:** a real counterfeiter might never see the detective’s reasoning. In a GAN, the Generator receives *mathematical feedback* (gradients) from the Discriminator, telling it precisely which direction to change its output. We’ll see this in section 12.\n\nHere is one complete cycle of the loop. Most pieces will look familiar from how any neural network trains:\n\nOne GAN training cycle\n\nRandom Noise → Generator → Fake Data → Discriminator\n\n → Real or Fake? → Calculate Loss → Backpropagation\n\n → Update Networks → Repeat\n\nIt’s the same machinery as any neural network: a **forward pass**, a **loss**, **backpropagation**, and **gradient descent**. The twist is that there are **two networks**, each with its own loss, each trying to beat the other.\n\nThe Generator doesn’t start from a picture. It starts from a **random vector**, called a **latent vector** or **noise vector**:\n\n[0.21, -0.73, 0.45, 0.91, …]\n\n ↓\n\n Generator\n\n ↓\n\n Generated Image\n\nIn my notes, this input is *sampled from a uniform distribution* and “forms the starting point of generating.” (Many implementations use a normal distribution instead. Either way, it’s just random numbers.)\n\n**Why random?** Because different random inputs lead to **different outputs**. The noise is the source of variety. If you gave the Generator the same input every time, it would produce the same image every time. The Generator’s real job is to learn a transformation that turns meaningless random numbers into **structured data**.\n\nOver training, the Generator learns the patterns that make real data look real. For images, it builds up from simple to complex, much like the CNN story told in reverse:\n\nPixels\n\n ↓\n\n Edges\n\n ↓\n\n Shapes\n\n ↓\n\n Textures\n\n ↓\n\n Objects\n\n ↓\n\n Realistic Image\n\nStructurally, a typical image Generator looks like this (straight from my notes): an **input** noise vector, **fully connected layers** that reshape the noise, **batch normalization**, **activation functions** (ReLU / LeakyReLU), **transposed convolutional layers**, then **reshaping and an output layer**.\n\nInside the Generator\n\nThe diagram shows a small MNIST-sized Generator:\n\n1. **100 random numbers** go in.\n\n2. A **Dense** layer expands them, followed by **BatchNorm** and **LeakyReLU**.\n\n3. **Reshape** turns the long vector into a small 7×7 feature map with 128 channels.\n\n4. **Transposed convolutions** grow that into 14×14, then 28×28: the size of an MNIST digit.\n\n5. The last layer uses **tanh**, so pixel values land between −1 and 1.\n\nIf a convolution *shrinks* an image into features (the way CNNs do), a **transposed convolution** does the opposite: it *grows* features back into an image.\n\nThe Discriminator learns to separate two groups:\n\nReal Data ──────→ Real\n\n \\\n\n → Discriminator\n\n /\n\n Generated Data ──→ Fake\n\nIt’s a CNN-style classifier. Per my notes, its typical structure is: **convolutional layers**, **batch normalization**, **LeakyReLU**, **pooling layers** (many modern designs use strided convolutions for downsampling instead), a **fully connected layer**, and an **output layer** giving a single value between 0 and 1.\n\nInside the Discriminator\n\n**Why LeakyReLU instead of ReLU?** My notes explain it from the CNN/activation chapter: LeakyReLU fixes the “dead neuron” problem (ReLU outputs zero for every negative input), keeps a small gradient for negative values, and gives *faster and more stable training*. That matters a lot in GANs, where stable gradients are hard to come by.\n\nThis is the heart of the idea:\n\nGenerator:\n\n “Can I fool you?”\n\n \n\n VS\n\n \n\n Discriminator:\n\n “Can I detect you?”\n\nDuring training:\n\n• The Generator improves at creating realistic data\n\n• The Discriminator improves at detecting generated data\n\n• Both networks keep adapting to each other\n\nNeither one is handed the “right answer” for what a good fake looks like. The Generator only learns because the Discriminator keeps pushing back.\n\nHere’s what’s unusual: the two networks are optimized for **opposite objectives.**\n\n**The Generator wants** the Discriminator to classify its generated data as *real*.\n\n**The Discriminator wants** to correctly classify *both* real and generated data.\n\nTwo players, opposite goals\n\nBoth use **binary cross-entropy** to calculate their loss (my notes: “both G and D use binary cross entropy to calculate loss”). It’s the same loss we use for any yes/no classifier.\n\nThe original paper combines both goals into one “minimax” objective:\n\nOne score, two players pulling in opposite directions. (In practice, implementations often tweak the Generator’s loss slightly so it learns better early on. The idea is the same: the Generator is rewarded when the Discriminator is fooled.)\n\n**Step 1:** Give random noise to the Generator\n\n**Step 2:** The Generator creates fake data\n\n**Step 3:** Give real and fake data to the Discriminator\n\n**Step 4:** The Discriminator predicts real or fake\n\n**Step 5:** Calculate both losses\n\n**Step 6:** Update the Discriminator\n\n**Step 7:** Update the Generator\n\n**Step 8:** Repeat\n\nIn steps 6 and 7 the two networks take turns. While one is being updated, the other is held still. (In the Keras code later, we’ll compute both gradients from the same forward pass and apply them together, which is a common simplification.)\n\nThe Generator never sees a real image as its target. So how does it learn?\n\nBad Image\n\n ↓\n\n Discriminator detects fake\n\n ↓\n\n Generator receives gradient\n\n ↓\n\n Generator updates weights\n\n ↓\n\n Better Image\n\n ↓\n\n Repeat\n\n**The Generator improves because the Discriminator provides a learning signal.**\n\nGradient flow\n\nDuring the **forward pass**, noise travels through the Generator, becomes a fake image, passes through the Discriminator, and produces a loss. During the **backward pass**, the gradient flows from the loss, *through the Discriminator*, and into the Generator. It tells the Generator how to change each pixel so the Discriminator’s score goes up.\n\nThat’s ordinary backpropagation and gradient descent, now passing through two networks chained together. The Discriminator is updated to catch fakes better; the Generator is updated to fool it.\n\nOnce training is done, you don’t need the Discriminator anymore. Generating is just:\n\nRandom Vector\n\n ↓\n\n Generator\n\n ↓\n\n Learned Features\n\n ↓\n\n Image\n\nRandom Noise\n\n ↓\n\n GAN\n\n ↓\n\n New Face\n\nEach new random vector gives a new output. The generated face is a **new synthetic example**, not a copy-paste of one training photo. It’s a fresh sample drawn from what the Generator learned about the data.\n\nMy notes list these applications:\n\n1. **Image generation**\n\n2. **Data augmentation** (more training examples when real data is limited)\n\n3. **Style transfer**\n\n4. **Super-resolution** (sharpening or enlarging images)\n\n5. **Generating art**\n\n6. **Image-to-image translation**\n\nOther uses include **synthetic datasets** and research in **synthetic media**.\n\nA note on responsible use: the same technology that makes artwork and helpful training data can also create **misleading synthetic media**. Labeling generated content and using it honestly matters.\n\nThere’s a whole family. Here’s the quick tour, based on my notes:\n\nThe GAN family\n\n**Vanilla GAN.** The simplest form, built using MLP layers. It’s the foundation of everything else. Known for training instability and mode collapse.\n\n**DCGAN (Deep Convolutional GAN).** Replaces fully connected layers with convolutional ones. It’s highly stable, produces high-quality, sharp images, and is a common choice for image generation.\n\n**Conditional GAN (cGAN).** Adds conditional variables like labels, attributes or text, so you can *command* the generator to produce a certain output.\n\nLabel = “Dog”\n\n +\n\n Random Noise\n\n ↓\n\n Generator\n\n ↓\n\n Dog Image\n\n**CycleGAN.** Image-to-image translation that **doesn’t need matching pairs** of training examples.\n\nHorse → Zebra\n\n**StyleGAN.** Produces photorealistic, high-resolution faces with hierarchical control over “style.” Developed by NVIDIA.\n\n16. GAN vs a traditional neural network\n\nPeople sometimes ask “which is better, a GAN or a CNN?” That’s a bit like asking whether a *recipe* is better than an *oven*.\n\n• **CNN** is an **architecture** commonly used to process images.\n\nGAN vs CNN\n\nFor image tasks, a GAN’s Generator and Discriminator can themselves be **CNN-based**. That’s exactly what DCGAN does.\n\nGANs are impressive and famously finicky. Here are the big three.\n\nThe Generator finds a few outputs that reliably fool the Discriminator and produces *only* those, instead of the full variety of the data.\n\nMode collapse\n\nExpected: Generated:\n\n Cat A Cat A\n\n Cat B Cat A\n\n Cat C Cat A\n\n Cat D Cat A\n\nOne mitigation from my notes: **minibatch** techniques, which let the Discriminator look at groups of samples rather than one at a time, so it can notice when everything looks the same.\n\nTraining two networks at once is hard: they’re two models of a neural network, and the optimization landscape is tricky. They can become unbalanced:\n\nTraining instability\n\n• **An overpowering Discriminator.** It gets so good that its strong output gives the Generator almost nothing useful to learn from. Helpful tricks: **label smoothing**, **feature matching**, or **slowing down** the Discriminator’s training.\n\n• **A weak Discriminator.** It can’t give the Generator meaningful feedback to improve.\n\nWhen things go wrong, the usual advice is to choose a GAN architecture that suits the problem and rethink the optimization setup.\n\nHow do you score a “good” face or digit? Measuring both the **quality** and the **diversity** of generated data isn’t straightforward, and often comes down to a mix of metrics and human judgment.\n\nMNIST (handwritten digits) is the ideal beginner dataset, because you can *see* whether the output looks like a digit.\n\nLoad Dataset → Build Generator → Build Discriminator → Generate Fake Images\n\n → Train Discriminator → Train Generator → Repeat\n\n**Setup and data**\n\nimport tensorflow as tf\n\n from tensorflow.keras import layers, models\n\n import matplotlib.pyplot as plt\n\n \n\n NOISE_DIM = 100\n\n BATCH_SIZE = 128\n\n \n\n (x_train, _), _ = tf.keras.datasets.mnist.load_data()\n\n x_train = x_train.astype(“float32”)\n\n x_train = (x_train — 127.5) / 127.5 # scale pixels to [-1, 1]\n\n x_train = x_train[…, None] # shape: (60000, 28, 28, 1)\n\n \n\n dataset = (\n\n tf.data.Dataset.from_tensor_slices(x_train)\n\n .shuffle(60000)\n\n .batch(BATCH_SIZE, drop_remainder=True)\n\n )\n\n**The Generator:** noise in, 28×28 image out.\n\ndef build_generator():\n\n return models.Sequential([\n\n layers.Input(shape=(NOISE_DIM,)),\n\n layers.Dense(7 * 7 * 128, use_bias=False),\n\n layers.BatchNormalization(),\n\n layers.LeakyReLU(0.2),\n\n layers.Reshape((7, 7, 128)),\n\n layers.Conv2DTranspose(64, 4, strides=2, padding=”same”, use_bias=False), # 14x14\n\n layers.BatchNormalization(),\n\n layers.LeakyReLU(0.2),\n\n layers.Conv2DTranspose(1, 4, strides=2, padding=”same”, activation=”tanh”), # 28x28\n\n ])\n\n**The Discriminator:** image in, one “realness” score out.\n\ndef build_discriminator():\n\n return models.Sequential([\n\n layers.Input(shape=(28, 28, 1)),\n\n layers.Conv2D(64, 4, strides=2, padding=”same”),\n\n layers.LeakyReLU(0.2),\n\n layers.Dropout(0.3),\n\n layers.Conv2D(128, 4, strides=2, padding=”same”),\n\n layers.BatchNormalization(),\n\n layers.LeakyReLU(0.2),\n\n layers.Dropout(0.3),\n\n layers.Flatten(),\n\n layers.Dense(1), # raw score (logit)\n\n ])\n\n \n\n generator = build_generator()\n\n discriminator = build_discriminator()\n\n**Losses and optimizers**\n\nbce = tf.keras.losses.BinaryCrossentropy(from_logits=True)\n\n \n\n def discriminator_loss(real_out, fake_out):\n\n real_loss = bce(tf.ones_like(real_out), real_out) # real should score 1\n\n fake_loss = bce(tf.zeros_like(fake_out), fake_out) # fake should score 0\n\n return real_loss + fake_loss\n\n \n\n def generator_loss(fake_out):\n\n return bce(tf.ones_like(fake_out), fake_out) # G wants fakes scored as real\n\n \n\n g_opt = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)\n\n d_opt = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)\n\n**One training step:** the 8-step cycle from earlier, in code.\n\n@tf.function\n\n def train_step(real_images):\n\n noise = tf.random.normal([BATCH_SIZE, NOISE_DIM])\n\n \n\n with tf.GradientTape() as g_tape, tf.GradientTape() as d_tape:\n\n fake_images = generator(noise, training=True) # steps 1–2\n\n \n\n real_out = discriminator(real_images, training=True) # steps 3–4\n\n fake_out = discriminator(fake_images, training=True)\n\n \n\n g_loss = generator_loss(fake_out) # step 5\n\n d_loss = discriminator_loss(real_out, fake_out)\n\n \n\n d_grads = d_tape.gradient(d_loss, discriminator.trainable_variables)\n\n g_grads = g_tape.gradient(g_loss, generator.trainable_variables)\n\n d_opt.apply_gradients(zip(d_grads, discriminator.trainable_variables)) # step 6\n\n g_opt.apply_gradients(zip(g_grads, generator.trainable_variables)) # step 7\n\n return g_loss, d_loss\n\n**Training loop that saves samples so you can watch it learn**\n\nseed = tf.random.normal([16, NOISE_DIM]) # same 16 noise vectors every time\n\n \n\n def save_samples(epoch):\n\n images = generator(seed, training=False)\n\n fig, axes = plt.subplots(4, 4, figsize=(4, 4))\n\n for ax, img in zip(axes.flat, images):\n\n ax.imshow(img[:, :, 0] * 0.5 + 0.5, cmap=”gray”) # back to [0, 1]\n\n ax.axis(“off”)\n\n plt.savefig(f”epoch_{epoch:03d}.png”, bbox_inches=”tight”)\n\n plt.close(fig)\n\n \n\n EPOCHS = 50\n\n for epoch in range(1, EPOCHS + 1):\n\n for batch in dataset:\n\n g_loss, d_loss = train_step(batch)\n\n print(f”Epoch {epoch}: G loss {g_loss:.3f} | D loss {d_loss:.3f}”)\n\n if epoch in (1, 10, 25, 50):\n\n save_samples(epoch)\n\nBecause we reuse the **same** seed every time, the 16 images in each saved grid are directly comparable: it’s the same 16 noise vectors being interpreted by an increasingly skilled Generator.\n\nThis is the most satisfying part of building a GAN: **watching the Generator improve.**\n\nEpoch 1 → Random-looking images\n\n Epoch 10 → Basic shapes\n\n Epoch 25 → Recognizable patterns\n\n Epoch 50 → More realistic samples\n\nEpoch template\n\n📌 **Your turn:** run the code above, then place your four saved grids (epoch_001.png, epoch_010.png, epoch_025.png, epoch_050.png) side by side where the template shows. The exact look depends on your run and settings, so use your own real outputs rather than the rough pattern above.\n\nEarly on you’ll see static-like blobs. Later, stroke-like shapes appear, and eventually many of the samples start to look like handwritten digits.\n\n1. **Thinking the Generator sees real images.** It never does; it only receives gradients through the Discriminator.\n\n2. **Letting the Discriminator win too easily.** If its loss drops to ~0 quickly, the Generator may stop learning.\n\n3. **Watching only the loss.** In GANs, a falling loss doesn’t guarantee better images. Look at the samples.\n\n4. **Forgetting to scale pixels** to match the Generator’s tanh output range (−1 to 1).\n\n5. **Expecting stable training.** Run-to-run variation is normal; tuning learning rates and architecture matters.\n\n6. **Treating “GAN” and “CNN” as alternatives.** One is a framework; the other an architecture.\n\n**What does GAN stand for?** Generative Adversarial Network.\n\n**Why does a GAN need two networks?** One network creates and the other judges. The judge’s feedback is what teaches the creator. Without an opponent, the Generator has no signal about what “realistic” means.\n\n**What is the Generator?** The network that turns random noise into synthetic data.\n\n**What is the Discriminator?** A binary classifier that decides whether a sample is real or generated.\n\n**What is random noise?** A random vector (the latent vector) that gives the Generator a starting point and is the source of variety in the outputs.\n\n**How does the Generator learn?** From gradients that flow back through the Discriminator. They show how to change the output to look more real.\n\n**How does the Discriminator learn?** Like any classifier: it’s shown real samples labeled “real” and generated samples labeled “fake,” and it minimizes its classification loss.\n\n**What is mode collapse?** When the Generator produces only a narrow range of outputs instead of the full variety in the data.\n\n**Can GANs create completely new data?** They create new samples that follow the patterns they’ve learned, and generally aren’t copies of individual training examples. But they’re limited by what they learned from: “new” here means new combinations within the learned patterns, not invention from nothing.\n\n**GAN vs CNN?** A CNN is an architecture. A GAN is a framework that can use CNNs inside both of its networks.\n\n**GAN vs traditional neural network?** Traditional networks map input to a prediction. A GAN pairs two networks to learn to generate realistic samples.\n\n**Aren’t autoencoders generative too?** A plain autoencoder compresses data into a latent space and reconstructs it, which is great for dimensionality reduction, denoising and anomaly detection, but it’s not built to generate new content. **VAEs (Variational Autoencoders)** go a step further: the encoder outputs the parameters of a probability distribution, we sample from it, and the decoder turns the sample into data. That makes them generative. It’s another road to the same destination, which we won’t go down today.\n\nRandom Noise\n\n ↓\n\n Generator\n\n ↓\n\n Fake Data\n\n ↓\n\n Discriminator ← Real Data\n\n ↓\n\n Feedback\n\n ↓\n\n Improve Generator\n\n ↓\n\n More Realistic Data\n\n• A GAN is **two networks**: a Generator (the forger) and a Discriminator (the detective).\n\n• The Generator turns **random noise** into data; the Discriminator is a **binary classifier**.\n\n• They have **opposite goals** and both use **binary cross-entropy**.\n\n• The Generator learns from **gradients that flow back through the Discriminator**.\n\n• Watch out for **mode collapse** and **training instability**.\n\n• GANs are a *framework*; they often use **CNNs** inside.\n\n**A GAN learns through competition: the Generator learns to create increasingly realistic data, while the Discriminator learns to distinguish generated data from real data.**\n\nGANs showed how neural networks can **generate** data by competing. Good next steps: train the MNIST GAN above and save your own epoch-by-epoch samples, try a **conditional GAN** so you can ask for a specific digit, and compare GANs with other generative approaches such as **VAEs (Variational Autoencoders)** and Transformer-based generators.\n\n*If this post helped, a clap or a follow means a lot, and I’d love to see your own epoch-by-epoch samples in the comments.*\n\n[GANs Explained: How AI Learns to Create New Data](https://pub.towardsai.net/gans-explained-how-ai-learns-to-create-new-data-9a8a124f77bd) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/gans-explained-how-ai-learns-to-create-new-data", "canonical_source": "https://pub.towardsai.net/gans-explained-how-ai-learns-to-create-new-data-9a8a124f77bd?source=rss----98111c9905da---4", "published_at": "2026-10-03 13:31:01+00:00", "updated_at": "2026-10-03 14:08:32.174025+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "neural-networks", "generative-ai", "ai-research"], "entities": ["Generative Adversarial Network", "Generator", "Discriminator", "Keras", "MNIST", "CNN", "RNN", "LSTM"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gans-explained-how-ai-learns-to-create-new-data", "markdown": "https://wpnews.pro/news/gans-explained-how-ai-learns-to-create-new-data.md", "text": "https://wpnews.pro/news/gans-explained-how-ai-learns-to-create-new-data.txt", "jsonld": "https://wpnews.pro/news/gans-explained-how-ai-learns-to-create-new-data.jsonld"}}