# Fix GPU Thermal Paste Pump-Out: A Repad Guide for 3090/4090 Rigs

> Source: <https://www.mindstudio.ai/blog/gpu-thermal-pad-maintenance-guide/>
> Published: 2026-09-12 00:00:00+00:00

# Fix GPU Thermal Paste Pump-Out: A Repad Guide for 3090/4090 Rigs

Diagnose GPU throttling from thermal paste pump-out and learn how to repad and repaste a GPU to lower temperatures on multi-GPU AI rigs.

## What is thermal paste pump-out and why does it hurt GPU performance?

Thermal paste pump-out happens when the liquid metal or paste between a GPU die and its cooler gets physically displaced over time from repeated heating and cooling cycles. The paste migrates outward, leaving a thin, dry patch in the center of the die where contact pressure is highest. The result is a GPU that shows normal-looking temperatures on paper (sometimes in the low 70s Celsius) but can’t hold its boost clock. Fans spin at 100% while the card keeps dropping into a lower performance state. That mismatch, high fan speed plus falling clocks despite “acceptable” temps, is the tell. It usually means the die isn’t transferring heat efficiently anymore, even though the sensor reading hasn’t spiked into thermal shutdown territory yet.

This matters more now than it used to. Anyone running local AI workloads, especially agentic or multi-step inference loads on cards like the RTX 3090 or 4090, is keeping GPUs under sustained load in ways gaming never really did. Sustained load accelerates pump-out. A card sitting at high utilization for hours or days at a time is a very different thermal stress test than short gaming sessions.

## TL;DR

- **Pump-out** shows up as pegged fans and clock throttling even when reported die temperatures look fine, and it’s a strong sign the thermal interface material has physically shifted off the die.
- Cards under **sustained AI workloads** for extended periods experience more thermal cycling stress than typical gaming use, making pump-out more likely to appear after a few years of heavy use.
- Diagnosing which physical card corresponds to which GPU ID requires running `nvidia-smi -L` , removing cards one at a time, and rebooting to match device order to a labeled sticker on the card.
- Full maintenance means measuring existing **thermal pads with digital calipers** rather than trusting published spec sheets, since pad thickness can vary by revision even within the same card model.
- **Thermal Grizzly’s Phase Sheet PTM** was used as a paste replacement material aimed at resisting pump-out longer than traditional paste, and chilling it briefly before application makes it easier to peel cleanly.
- Careful documentation, photos of screw placement, pad locations, and pad dimensions, prevents reassembly mistakes, since GPU teardown involves many differently sized screws and pad thicknesses.
- Setting a **power limit** on GPUs (even reducing power by roughly 20%) can meaningfully cut total system wattage across a multi-GPU rig with minimal performance loss, which also reduces thermal stress going forward.

## Other agents start typing. Remy starts asking.

Scoping, trade-offs, edge cases — the real work. Before a line of code.

## How do you know if your GPU needs a repaste or repad?

The clearest sign is a boost clock that won’t hold steady combined with fans running flat out. If a card is bouncing down into a lower performance state (visible in monitoring tools as a P-state drop) while temperatures sit in a range that doesn’t look alarming, that’s the classic pattern of pump-out rather than a straightforward overheating problem. A truly overheating card usually shows high temps and high fans together. Pump-out shows fans maxed with only moderate temps, because the sensor is reading a hot spot condition, not a globally hot die, and the card’s firmware throttles preemptively to protect itself.

Opening the card confirms it. Pump-out looks distinct once you remove the cooler: a dry patch in the center of the die where the heat sink made contact, with the more viscous remaining material pushed out to the edges. Visually, it looks like the paste “blew out” from the middle, leaving almost nothing where contact pressure is greatest and the most cooling is needed.

## How do you disassemble a GPU without ruining it?

The teardown process is straightforward but unforgiving of sloppy notes. Before touching any hardware, power supplies should be fully switched off, unplugged, and drained by holding the power button for several seconds to clear residual charge.

A few practices make disassembly safe to reverse:

Label every GPU before removal. Running `nvidia-smi -L` on Windows or Linux lists device IDs, but it won’t tell you which physical slot corresponds to which listed device. The reliable method is removing cards one at a time, rebooting, and re-running the command to match device order to a physical card by elimination. Once identified, a label maker sticker on the card’s edge avoids repeating this process during every future maintenance cycle.

Use separate trays for screws from different assembly stages (fan shroud, backplate, faceplate), since screw lengths often differ by location and mixing them up can crack a PCB or leave gaps in cooler contact. Photographing screw layouts and pad locations with a phone before removing anything is cheap insurance against forgetting where a nonstandard screw or oddly shaped pad belongs.

Pads themselves should be documented individually: photograph the layout, then as each pad comes off, note its dimensions in a photo editor or notes app before moving to the next. Storing removed pads in a plastic bag keeps them viable for reuse if only some need replacing.

## Why do you need to measure your own thermal pads instead of trusting online specs?

Because published pad thickness guides for specific card models can simply be wrong. Manufacturing revisions change pad depths even within the same card model line, and trusting a measurement found online rather than checking the physical card can lead to ordering the wrong thickness and wasting money on unusable pads.

### Everyone else built a construction worker.

We built the contractor.

One file at a time.

UI, API, database, deploy.

The fix is measuring with digital calipers directly. Zero the calipers, then measure a removed pad’s thickness by closing the caliper jaws on it. Getting an accurate reading sometimes takes a couple of tries (checking against a dark background can help visibility), narrowing down from a rough guess to a precise figure like 1.5mm rather than assuming 1mm or 2mm. Pad length and width should be measured and recorded the same way. Most cards need a mix of thicknesses (commonly somewhere across the 1mm to 3mm range depending on location) rather than a single uniform pad, since different components on the board sit at different heights relative to the heat sink.

For anyone maintaining several GPUs at once, buying pad material in bulk sheets rather than pre-cut, model-specific kits can be more economical, provided the measurements are accurate.

## What is Thermal Grizzly Phase Sheet PTM and how is it applied?

Phase Sheet PTM is a solid-at-room-temperature thermal interface material from Thermal Grizzly designed as an alternative to conventional paste, intended to resist migrating off the die the way liquid paste can over years of thermal cycling. It ships as a sheet that’s cut to size for the die.

Application benefits from chilling the material first, briefly in a freezer or longer in a refrigerator, which firms it up and makes it easier to peel cleanly off its backing without tearing. Before applying it, the die needs a thorough cleanup with isopropyl alcohol at 91% concentration or higher (lower concentrations like 70% leave residue behind as they evaporate). Cleaning should avoid saturating nearby thermal pads with alcohol, since that can degrade them. Working the alcohol across the die in multiple directions, along the grain of any residue and against it, helps lift stubborn remnants of the old pump-out debris before the new phase-change material goes on.

## Does lowering GPU power limits actually help long-term reliability?

Reducing a GPU’s power limit is a low-effort way to cut heat generation and, by extension, reduce the thermal cycling that contributes to pump-out over years of use. Dropping power by around 20% typically produces a fairly small performance hit relative to the wattage saved, since GPUs (and 3090s in particular) often run past their efficient point on the voltage/frequency curve at stock settings. Across a rig with four GPUs, that kind of reduction is enough to meaningfully lower total system draw, shifting a setup from consuming somewhere around 1.3 to 1.4 kilowatts down closer to 1 kilowatt under full load. For anyone running multiple cards continuously for AI workloads, that’s a real difference in electricity cost, heat output in the room, and power supply headroom, not just a marginal tweak.

## Frequently Asked Questions

### How often should GPU thermal paste or pads be replaced?

The transcript doesn’t give a universal number, but it points to roughly a multi-year window (on the order of several years of heavy use) as a serious consideration point, especially for high-heat cards like the 3090 running sustained workloads. Cards under continuous AI inference load are likely to need attention sooner than cards used intermittently for gaming.

### What tools do you need to repad or repaste a GPU?

A small magnetic-tip screwdriver, a mag tray (or multiple trays) to keep screws sorted by disassembly stage, digital calipers to measure pad thickness and dimensions, isopropyl alcohol at 91% or higher, cotton swabs, a phone camera for documentation, and the replacement thermal interface material and pads sized to the measurements taken from the card itself.

### Can you tell which GPU is which without opening the case?

## Other agents ship a demo. Remy ships an app.

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

Yes, using `nvidia-smi -L` to list device IDs, then removing cards one at a time and rebooting to determine physical position by elimination. Once matched, labeling each card physically avoids repeating this process during future maintenance.

### Is Thermal Grizzly Phase Sheet PTM better than regular thermal paste?

It’s presented as an option built specifically to resist the pump-out failure mode that affects standard paste over time, making it a candidate for anyone doing maintenance on GPUs expected to run under sustained load for years. No independent benchmark figures were provided in the source material, so comparative performance numbers should be checked against the manufacturer’s own documentation.

### Does undervolting or power-limiting a GPU hurt AI inference performance?

Based on the source, a power limit reduction of around 20% produces a fairly small performance impact relative to the wattage saved, making it a reasonable default for multi-GPU rigs where total power draw and heat matter as much as peak throughput.
