# What is the DLSS5 model architecture?

> Source: <https://gist.github.com/madebyollin/55c703a34bf90962844edcd68d04e32e>
> Published: 2026-08-30 15:12:05+00:00

Let's attempt to figure out the model architecture from DLSS5's binary! See [https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee](https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee) for context.

Warning

The remainder of this gist was authored by Codex 5.6

The strongest current model is a **heterogeneous, hierarchical, recurrent,
single-pass neural renderer**. Its visible body combines tiled 1H/2H/4H/8H
Swin families, a compound 16H split-Swin core, 2D and 1D ViT compound blocks,
specialized 1x1 width transitions, and a decoder with upsample/skip routing.

The most important difference from a conventional diffusion implementation is that the runtime exposes no timestep, noise schedule, or loop over denoising steps. However, the first 1H pre-block internally synthesizes Gaussian-like values from a seed-like scalar and tile coordinates, and those values enter the first tensor-core computation. The best inference-level classification is therefore one-pass latent-conditioned generative rendering.

```
flowchart LR
  subgraph raw[Engine and temporal resources]
    color["Color / RGB"]
    mv["MVec + scale"]
    depth["Depth"]
    prev["dlssnr_prev_output"]
    control["ControlMask / UI / Backbuffer"]
  end

  submit["CG2R network-manager submission\nresource packing + CUBIN bindings"]
  pre["1H pre-block / input adapter\nRGB texture + feature packing"]
  rng["Internal Gaussian path\nseed-like field + tile coordinates\n3 generated FP16 lanes + 1.0"]
  e1["1H / 32\ntiled Swin"]
  e2["2H / 64\ntiled Swin"]
  e3["4H / 128\ntiled Swin"]
  e4["8H / 256\ntiled Swin"]
  core["16H split-Swin compound core\nFFN + fused QKVAttn + projections/pool\nfinal head: 512 -> 1024"]
  bridge["CCDecInputUpsample\n1024 -> 512\nmain + skip\noutput + reduce/overlap/conv-tap CBs"]
  d4["8H decoder\nupsample + skip"]
  d3["4H decoder\nupsample + skip"]
  d2["2H decoder\nupsample + skip"]
  d1["1H decoder\nupsample + skip"]
  post["1H post-block\n2 inputs -> CudaSurface\nmask / simple-blend variants"]
  output["Output\nplus history update"]

  subgraph islands["Registered ViT islands; active placement unresolved"]
    vit2d["ViT 2D compound block\nFFN + QKV + attention + projection"]
    vit1d["ViT 1D compound block\nFFN + QKV + attention + projection"]
    repack["2D <-> 1D repack"]
    vit2d -.-> vit1d -.-> repack
  end

  color --> submit
  mv --> submit
  depth --> submit
  prev --> submit
  submit --> pre
  rng -. generated inside pre-block .-> pre
  pre --> e1 --> e2 --> e3 --> e4 --> core --> bridge
  e4 -. skip .-> d4
  e3 -. skip .-> d3
  e2 -. skip .-> d2
  e1 -. skip .-> d1
  bridge --> d4 --> d3 --> d2 --> d1 --> post --> output
  control --> post
  output -. writes .-> prev

  classDef observed fill:#d9f2e6,stroke:#24734d,color:#123b29;
  classDef inferred fill:#fff1cc,stroke:#9a6b00,color:#4d3500;
  class submit,pre,rng,e1,e2,e3,e4,core,bridge,d4,d3,d2,d1,post observed;
  class raw,prev,output,vit2d,vit1d,repack inferred;
```

The solid path is a family-level skeleton, not a claim that each displayed
level maps to exactly one serialized `blockN`

. The weight resource contains
many repeated records and sub-block tensors. The dashed skip edges are
structurally supported by the decoder kernels and wrapper contracts, but the
exact merge operation is not yet separated from the fused code.

The ViT/1D blocks are shown as a separate registered island because their internal inventories are clear while the active descriptor sequence does not yet establish whether they sit before the 16H core, between the core and decoder, or in a parallel conditioning branch.

| property | recovered value | confidence |
|---|---|---|
| target DLL | NVIDIA NGX `rel_310_8` , `nvngx_dlssnr.dll` |
high |
| SHA-256 | `e16bcf15e16e13f527491cdf7845b2fe6521a738d8f7c9c721866a8496e1fc8e` |
high |
| learned-resource payload | `WEIGHTS_HT/1033` , 147,695,410 bytes |
high |
| resource records | exactly 153 parsed records across 71 logical block IDs (`block0` –`block70` ) |
high |
| embedded CUDA | 15 fatbins and 231 public-inventory kernel names | high |
| GPU target | `sm_120` / Blackwell only in this build |
high |
| architecture | hierarchical tiled Swin + split-Swin/ViT/1D compound runtime + decoder routing | high for vocabulary; medium for exact topology |
| visible hierarchy | 1H/32 → 2H/64 → 4H/128 → 8H/256, with a 16H/512 central core | high for families; medium for placement/counts |
| central compound block | six specialized layer types, including fused QKV-attention and a final width expansion | high |
| decoder bridge | `CCDecInputUpsample` , 1024 → 512, main + skip |
high |
| parameter count | approximately 148 million bytes of opaque learned-resource payload; exact independent scalar count remains unresolved | medium |
| temporal path | previous output bound into the network-manager submission and updated after evaluation | high structurally; medium for exact slot |
| internal random path | hash/coordinate-derived Gaussian-like FP16 lanes in the 1H pre-block | high for mechanics; medium for semantics |
| inference style | one forward submission; no visible timestep or denoising loop | high |

The public write-up's estimate of roughly 148 million FP8 parameters is
credible as a rounded model-size figure, but the static resource parser now
gives a more precise statement of what is actually measured. The
`WEIGHTS_HT/1033`

resource is 147,695,410 bytes and contains exactly 153 named
records. Their opaque payloads total 147,683,778 bytes; the remaining 11,632
bytes are the resource header, record metadata, padding, and footer. The
record tags are raw storage markers (including several non-printable-looking
values), not yet decoded E4M3/E5M3 labels. Some records also contain 2-byte or
small auxiliary payloads. Consequently, the payload total is not yet an exact
count of independent learned scalars.

What *is* now exact at the resource level is the named-record and block
inventory: 71 logical IDs, with a near-symmetric 4/5/8-layer core pattern.
That inventory materially constrains the topology even though the active
descriptor order and storage decoding remain separate questions.

The resource parser follows the record framing and validates the resource-size
field against the actual 147,695,410-byte blob. It recovers the public
write-up's 153-record count, including the small `block70.layer0.blend_scale`

record that is easy to miss if only `blockN.layerM.layer`

names are counted.
The resource contains 71 numeric block IDs (`block0`

through `block70`

). The
most informative aggregate pattern is:

| block IDs | records per block | payload bytes per block | structural reading |
|---|---|---|---|
`block0` –`block4` |
1 | 20–23 KiB | edge/adapter-size records |
`block5` –`block8` |
1 | 62–70 KiB | small transition records |
`block9` –`block14` |
1 | 197–230 KiB | medium transition records |
`block15` –`block22` |
1 | 689–820 KiB | wide single-layer records |
`block23` –`block29` |
4 | 1,968,192 | repeated 4-layer compound group |
`block30` |
5 | 2,492,496 | 5-layer compound boundary group |
`block31` –`block38` |
5 | 12,587,154 | eight repeated central 5-layer groups |
`block39` |
1 | 525,312 | central transition record |
`block40` –`block47` |
4 | 1,968,192 | repeated 4-layer compound group |
`block48` –`block55` |
1 | 689–821 KiB | wide single-layer records |
`block56` –`block61` |
1 | 197–230 KiB | medium transition records |
`block62` –`block69` |
1 | 21–70 KiB | small transition/adapter records |
`block70` |
2 | 21,810 | 2-byte `blend_scale` plus one layer record |

The ranges above intentionally summarize similar payload sizes rather than assigning each ID to an encoder or decoder stage. The center's eight repeated five-record groups, flanked by repeated four-record groups, are strong structural corroboration for a large compound core surrounded by hierarchical transitions. They do not, by themselves, prove the runtime's active execution order or map one serialized block to one displayed Swin family.

The feature wrapper exposes `DLSSNR.Color`

, `MVec`

, and `Depth`

as core inputs,
with optional `ControlMask`

, `UI`

, `UIAlpha`

, `Backbuffer`

, and
`BidirectionalDistortionField`

resources. It also exposes controls including
`Intensity`

, `LocalToneStrength`

, `LocalStructureStrength`

,
`SkinStructureStrength`

, `UseAutoMask`

, `UICorrection`

, `Enabled`

, `Reset`

,
and `DepthInverted`

.

The runtime maintains separate `dlssnr_original_color`

and
`dlssnr_network_output_scratch`

surfaces in addition to `dlssnr_prev_output`

.
The evaluator creates the scratch/original-color resources and invokes
`CG2RNetworkManager::Evaluate`

(`sub_180021bb0`

). The manager reads the
feature's previous-output field and sends it through the CUBIN resource-binding
path. A later copy/update path writes the persistent history resource.

This is strong evidence for recurrent history feedback into the learned network submission rather than a history buffer used only by final blending. The exact history representation, reprojection, and insertion point remain open: the current evidence does not prove whether history is concatenated at the pre-block, injected deeper in the hierarchy, or transformed by a dedicated temporal kernel.

The CPU runtime registers 1H, 2H, 4H, and 8H fused Swin families. Representative CUDA modules use tensor-core MMA, explicit shared-memory staging, and variants for input/output views, chaining, waiting, tile synchronization, upsampling, and down/scale routing. The family names encode widths and head counts:

| family | visible width/head clue | representative role |
|---|---|---|
`cc_tinlayout_fused_pre_block_swin_1h_32*` |
1H / 32-wide | RGB-aware pre-block and first fused block |
`cc_tinlayout_fused_swin_1h_32*` |
1H / 32-wide | regular fine-resolution local block |
`cc_tinlayout_fused_swin_2h_64_2*` |
2H / 64-wide | hierarchical local block |
`cc_tinlayout_fused_swin_4h_128_4*` |
4H / 128-wide | hierarchical local block |
`cc_tinlayout_fused_swin_8h_256_8*` |
8H / 256-wide | hierarchical local block and decoder counterpart |

The pre-block's template includes `FusedSwin2d1HConfig<32,...>`

and
`MpCubicSiluActivation`

. The regular 1H body has 512 MMA operations and
repeated `rsqrt`

operations; the representative 2H, 4H, and 8H bodies have
approximately 512, 448, and 576 MMA operations respectively. Non-FP8 shared
footprints grow from roughly 9 KiB to 17 KiB to 34 KiB across those wider
families, consistent with larger local tiles and more parallel channels.

The decoder-specific variants expose `_upsample`

names and the host runtime
has an explicit `cc_tinlayout_upsample_skip_block`

factory. This supports a
U-Net-like hierarchy with encoder features retained for decoder-side routing,
although the exact concat/add/gated merge is unresolved.

The host wrapper `CCSplitSwin16HBlock`

(`sub_180040260`

) recognizes six
specialized layer types:

```
CCSplitSwin16HFfwd
CCSplitSwin16HFfwdProj
CCSplitSwin16HQKVAttn
CCSplitSwin16HProj
CCSplitSwin16HProjPool
CCSplitSwin16HFinalHead
```

The corresponding CUDA family contains separate FFN, FFN-projection, QKV,
projection, projection+pool, and final-head entries. The host type is
explicitly `QKVAttn`

, and there is no standalone split-Swin attention entry;
the QKV family owns the fused middle operation. The base FFN entry contains
256 MMA operations, while the QKV entry contains 224 MMA operations plus
warp-level reduction, normalization, and reciprocal work.

The final-head entry is a central feature transition, not the final image output head. Its template is:

```
Conv2d1x1Config<1024,512,...>
```

The first convolution template channel is the output width, as independently
confirmed by the ViT FFN expand/contract pair. Therefore this layer expands
512 → 1024. The later decoder bridge uses `Conv2d1x1Config<512,1024,...>`

and
contracts 1024 → 512. The host `CCDecInputUpsample`

constructor independently
checks `0x400 → 0x200`

and identifies its specialization as “1024->512
(dec5)”.

The host-side layer contracts make the split-Swin interface more specific than
the kernel names alone. `CCSplitSwin16HFfwd`

is a one-input/one-output layer;
`CCSplitSwin16HFfwdProj`

takes `(skip, src)`

and produces one output;
`CCSplitSwin16HQKVAttn`

is one input/one output; `CCSplitSwin16HProj`

takes
`(src, skip)`

and produces one output; `CCSplitSwin16HProjPool`

takes two
inputs and produces two outputs; and `CCSplitSwin16HFinalHead`

is one
input/one output. This is consistent with an internally routed compound
block, rather than six independent top-level network stages.

The 2D ViT wrapper (`sub_180041ca0`

) recognizes five layer types:

```
CCVitFfnExpand
CCVitFfnContract
CCVitQKV
CCVitAttention
CCVitProjection
```

The 1D wrapper (`sub_180043260`

) recognizes corresponding 1D variants and
explicitly checks for five layer descriptors. The CUDA family adds
`cc_vit_1d_repack_2d_to_1d`

and `cc_vit_1d_repack_1d_to_2d`

entries.

Representative ViT QKV entries contain 192 MMA operations, consistent with three projections. Separate attention entries contain 128 MMA operations for QK-like work and weighted V accumulation. FFN expand/contract entries are separate; the 2D expand/contract templates are 1024 → 4096 and 4096 → 1024.

These blocks are high-confidence runtime capabilities and compound structures, but their active placement remains unresolved because the generic factory and constructor paths expose type dispatch rather than the serialized active network sequence.

The following contracts come from the host wrappers' validation strings and are more reliable than inferring interfaces from register allocation alone:

| layer family | inputs | outputs / side buffers | confidence |
|---|---|---|---|
| 1H pre-block | RGB texture input | one output; `_ds` exposes `(pool, swin)` |
high |
| 1H/2H/4H/8H regular Swin | one tensor; upsample variant additionally takes `skip` |
one output; `_ds` exposes `(pool, swin)` |
high |
| split-Swin FFN / FFN-projection | one tensor / `(skip, src)` |
one output | high |
| split-Swin QKV-attention | one tensor | one output | high |
| split-Swin projection | `(src, skip)` |
one output | high |
| split-Swin projection+pool | two tensors | two outputs, likely pool plus transformed projection | high for arity; medium for identity |
| split-Swin final head | one tensor | one tensor, 512 → 1024 | high |
| 2D ViT FFN expand | one tensor | one output | high |
| 2D ViT FFN contract | `(skip, src)` |
`dst` plus reduce CB |
high |
| 2D ViT QKV | one tensor | `qry` , `key` , `val` plus reduce CB |
high |
| 2D ViT attention | `qry` , `key` , `val` |
one output | high |
| 2D ViT projection | `(src, skip)` |
`dst` plus reduce CB |
high |
| 1D ViT FFN/QKV/attention/projection | explicit token tensors and input/output CBs | explicit Q/K/V and destination/reduction CBs | high |
| decoder input upsample | `(main, skip)` |
`{tensor, reduce_cb, overlap_cb, conv_tap}` |
high |
| post-block | `(main, enc0 skip)` |
one `CudaSurface` output |
high |

The ViT QKV contract is especially informative: the runtime really does
materialize three logical projections (`qry`

, `key`

, `val`

) plus a reduction
control buffer at the wrapper boundary, even though the GPU implementation
fuses parts of their processing. The 1D wrapper likewise names explicit
input/output CBs for the token path and exposes 2D↔1D repack kernels. This
narrows the possible model interfaces substantially without claiming that
these ViT islands are active in the shipping forward path.

The specialized `CCDecInputUpsample`

wrapper checks the channel transition
`1024 → 512`

and requires two inputs named `(main, skip)`

. It also validates a
four-part output bundle:

```
{ tensor, reduce_cb, overlap_cb, conv_tap }
```

The matching normal decoder kernel is a no-activation
`Conv2d1x1Config<512,1024,...>`

implementation with 64 MMA operations, async
shared staging, and repeated vectorized output stores. The FP8 variant changes
the storage tile configuration but preserves the same learned 1x1 bridge
role. The PTX shows global reduction updates and a final integer release/store
path indexed by tile/block coordinates, which explains why the wrapper exposes
reduction and overlap control buffers in addition to the tensor result. The
exact parameter-slot identity of `reduce_cb`

, `overlap_cb`

, and `conv_tap`

is
not fully proven because the synchronization variants add pointer arguments,
but the four-output contract itself is explicit.

The post kernel family is likewise more than a generic image store. The base
`CCTinlayoutFusedPostBlockSwin1HLayer`

wrapper requires `(main, enc0 skip)`

and
one `CudaSurface`

output. Its kernel variants include `simple_blend`

,
`control_mask`

, and `full_rect`

forms, with corresponding FP8 variants. The
PTX has a late texture-sampling/control path and a four-component surface
write; the most defensible interpretation is three transformed output lanes
plus a fourth alpha/control-like lane, followed by mode-dependent blending.
The exact color-space transform and history-update equation remain open.

The CUDA templates name the activation `MpCubicSiluActivation`

. The emitted
code does not call a conventional transcendental sigmoid. Instead, the
representative low-precision path clamps the input and evaluates a cubic-like
piecewise approximation:

```
t = clamp(x, -4, +4)
p = (-0.0559082) * abs(t) + 0.447266
a = t * p + 0.894531
y = x * a
```

The constants are rounded to half precision in the recovered path. This is a hardware-friendly SiLU-like nonlinearity that keeps the fused blocks tensor- core oriented.

The representative hierarchical Swin and split-QKV kernels contain FP32
reductions followed by `rsqrt.approx.ftz.f32`

, gain/bias-like multiplies, and
reciprocal operations. This establishes an explicit variance- or energy-style
normalization stage. The binary evidence available so far does not prove
whether every instance subtracts a mean like LayerNorm or instead uses an
RMS/L2-style normalization, so the exact normalization equation remains open.

The 2D and 1D ViT attention families are not simple unfused matrix-multiply
wrappers. A representative `cc_vit_1d_attention`

path performs QK-like MMA,
applies a half2 affine/clamp transform, constructs a fast positive
exponential-like surrogate using bit manipulation, clamps the denominator to
approximately `6.2e-5`

, and accumulates V with reciprocal lanes. It is best
described as softmax-like attention; the recovered code does not expose a
canonical library softmax or `ex2.approx`

sequence.

The central split-Swin QKV family similarly fuses projection, reduction,
normalization, reciprocal, and middle-operation work. Because the host type is
`CCSplitSwin16HQKVAttn`

and there is no separate split-attention registration,
the exact split attention/mixer equation cannot yet be isolated from its QKV
kernel.

The embedded code targets Blackwell `sm_120`

and includes tensor-core FP8
conversion instructions such as `UF2FP.SATFINITE.E4M3.F16`

and
`F2FP.SATFINITE.E4M3.F16`

; both E4M3 and E5M3 strings are present. The `_fp8`

variants are the explicit low-precision weight path. Other suffixes encode
execution variants rather than new neural layers:

| suffix/family | likely role |
|---|---|
`_fp8` |
FP8 weight/dequantization path |
`_chained` , `_wait` , `_tilesync` |
dependency and tile-scheduling variants |
`_inpview` , `_outview` |
input/output view-layout variants |
`_upsample` , `*_ds` |
hierarchy transition or decoder variants |
`repack_2d_to_1d` , `repack_1d_to_2d` |
ViT 2D/1D layout conversion |
`control_mask` , `simple_blend` , `full_rect` |
post-block control/output modes |

This naming matrix is one reason the public kernel count is much larger than the number of conceptual neural layers.

The strongest generative clue is inside the 1H pre-block rather than in a
caller-visible input surface. The raw PTX for the representative pre entry
contains a 264-byte parameter block, with a scalar at offset `+200`

and
dimension-like values at `+208`

. The scalar is multiplied by
`-1640531527`

, then combined with tile coordinates in a hash-like sequence.

The resulting code consumes four uniform-like values and executes the following recognizable pattern:

- transform uniform values with
`lg2`

,`ln2`

,`-2`

, and`sqrt`

; - use
`sin`

and`cos`

with a`2π`

-like constant, as in a Box–Muller-style Gaussian construction; - convert generated values to FP16;
- combine three generated lanes, a constant
`1.0`

, and four sampled texture lanes; - stage the resulting vector through shared memory before the first MMA.

The normal, FP8, and downsampled pre variants share this signature. No caller-provided noise surface, timestep, noise schedule, or loop over denoising steps was recovered. The safest interpretation is therefore an internal deterministic coordinate/seed-derived Gaussian-like conditioning path. Its exact semantic role—latent injection, stochastic feature augmentation, or a specialized preconditioning operation—remains unresolved.

The inference binary strongly supports a single-pass latent-conditioned generative renderer, but it does not contain the training metadata needed to name the objective. The current evidence ranks the interpretations as follows:

| interpretation | assessment from the binary |
|---|---|
| conventional multi-step diffusion sampler | unlikely: no timestep/schedule state or denoising loop is visible in the analyzed path |
| one-step distilled diffusion, consistency, flow/meanflow, or DMD-like model | strongly compatible with the one-pass latent-conditioned body and internal Gaussian-like path |
| GAN-style single-pass generator | also compatible with the observed inference graph |
| deterministic neural renderer with a random-looking internal conditioner | cannot be ruled out from implementation alone |

DLSS5 appears to combine several performance choices that are mutually reinforcing:

- Blackwell-only
`sm_120`

kernels use tensor-core MMA paths and explicit FP8 conversion variants. - Windowed/tiled Swin-style operators keep attention local and fuse nearby projection, normalization, activation, and layout work.
- The 16H compound core specializes its FFN, QKV-attention, projection, pool, and feature-transition stages instead of dispatching generic primitives.
- Async dependency variants (
`_chained`

,`_wait`

,`_tilesync`

) let the runtime match kernel scheduling to the tile graph. - A single analyzed forward path avoids the cost of a conventional denoising loop.
- Recurrent history feedback supplies temporal context without requiring a large sliding window of prior frames.

The current architecture is strong enough for a report, but several details would require live execution traces, more host-side state recovery, or weight interpretation:

- the serialized active layer sequence, repetition counts, and exact placement of the 2D/1D ViT islands;
- the meaning and lifetime of the
`+200`

pre-block scalar and the precise mapping from generated lanes to network channels; - the pre-block's complete RGB/motion/depth packing and the exact shape of its internal Gaussian-like injection;
- the exact mean-subtracting-vs-RMS normalization equation;
- the full split-Swin QKV-attention/mixer equation;
- whether decoder skip routing is add, concatenate, gated blend, or a fused
equivalent, and the exact identity of the decoder
`reduce_cb`

,`overlap_cb`

, and`conv_tap`

pointer slots; - the precise post-block color transform and output/history update semantics for each mask and blend mode;
- the active descriptor's exact block sequence and the placement/repetition of the ViT/1D islands. The builder selects a runtime descriptor and dispatches through factories, but the serialized descriptor contents are not exposed as a simple string list in this build;
- the training provenance: one-step diffusion distillation, consistency/flow matching, DMD-like regression, GAN, or another objective.

The following host-side procedures were the most useful anchors in Hopper:

| location | evidence recovered |
|---|---|
`sub_18003c280` |
family registration for pre/post 1H, hierarchical Swin, split-Swin, ViT, decoder bridge, and clear callback |
`sub_1800326c0` |
weight binding names: `input_adapter_weight` , `weight1` , `weight2` , `ffn_cos_skip` , `qkv_weight` , `attn_scale` , `attn_bias` , `projection_weight` , `attn_cos_skip` |
`sub_18003cc80` |
pre-block dispatch and the explicit “pre-block requires rgb input” guard |
`sub_18003fdc0` |
single-layer dispatcher for 1H/2H/4H/8H Swin, fused pre/post blocks, and decoder input upsample |
`sub_180040260` |
six-layer `CCSplitSwin16HBlock` compound wrapper |
`sub_180041ca0` |
five-layer 2D `CCVitBlock` wrapper |
`sub_180043260` |
five-layer 1D `CCVit1DBlock` wrapper and repack expectations |
`sub_180073bf0` |
decoder bridge dimension checks and “1024->512 (dec5)” specialization |
`sub_18001f570` range |
active-network builder: descriptor selection, runtime dimension logging, consolidated aligned weight-heap sizing, and factory-driven network construction |
`sub_180021bb0` |
`CG2RNetworkManager::Evaluate` , resource setup, previous-output handling, and CUBIN binding path |
`tools/parse_dlss5_weight_map.py` |
validated 153-record `WEIGHTS_HT` map, 71 logical block IDs, payload totals, and raw storage tags |

The CUDA-side evidence comes from the 15 embedded `sm_120`

fatbins, the
representative PTX/SASS reductions, and the kernel-family naming matrix.
Static analysis establishes capabilities and likely dataflow; it does not by
itself prove the runtime's active descriptor sequence.

| aspect | DLSS 4.5 report | DLSS5 finding |
|---|---|---|
| core organization | compact local-transformer/U-Net-like body | heterogeneous hierarchical Swin + 16H split-Swin core + registered ViT islands |
| attention/mixing | local softmax attention followed by a softmax-free local mixer | fused local Swin families, split QKV-attention, and ViT softmax-like attention |
| activation | cubic SiLU approximation | same named `MpCubicSiluActivation` family is present |
| normalization | more fully characterized from the smaller graph | rsqrt/gain-style normalization recovered; exact mean-vs-RMS form remains open |
| temporal path | previous output/history participates in the graph | recurrent previous-output/history path is again visible, with no evidence of a large sliding window |
| output path | dynamic anisotropic Gaussian reconstruction filter was recovered | post-block mask/simple-blend/full-rect families are present, but a comparable output filter has not been recovered |
| latent/random path | no analogous internal Gaussian signature reported | pre-block contains an internal Gaussian-like, coordinate/seed-derived path |
| inference interpretation | deterministic neural upscaler with specialized local mixing | single-pass latent-conditioned generative neural renderer |
| training claim | architecture analysis did not require a training-objective claim | one-step distilled diffusion is plausible, but not distinguishable from GAN/flow/consistency alternatives from the DLL alone |

The comparison is intentionally asymmetric: DLSS4.5's smaller model was amenable to a more complete graph reconstruction, while the DLSS5 artifact exposes a larger, more heterogeneous runtime with several compound islands whose active ordering is not serialized in the public evidence we have.

DLSS5 is best described as a recurrent, single-pass, latent-conditioned generative neural renderer. Its body combines a 1H input adapter, hierarchical 1H/2H/4H/8H tiled Swin stages, a 16H split-Swin compound core with a 512 → 1024 feature expansion, a 1024 → 512 decoder bridge, decoder-side upsample/ skip blocks, and a 1H mask/blend post path. ViT and 1D transformer blocks are also registered as specialized compound islands, although their active placement is not yet proven.

The internal coordinate/seed-derived Gaussian-like path and the absence of a runtime timestep or denoising loop make one-step generative inference a strong architectural interpretation. They do not, however, identify whether NVIDIA trained it with meanflow, consistency distillation, DMD, GAN losses, or some other teacher/student objective. That distinction belongs in the “likely training neighborhood,” not in the list of facts directly recovered from the binary.
