Let's attempt to figure out the model architecture from DLSS5's binary! See https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee for context.
Warning
The remainder of this gist was authored by Codex 5.6
The strongest current model is a heterogeneous, hierarchical, recurrent, single-pass neural renderer. Its visible body combines tiled 1H/2H/4H/8H Swin families, a compound 16H split-Swin core, 2D and 1D ViT compound blocks, specialized 1x1 width transitions, and a decoder with upsample/skip routing.
The most important difference from a conventional diffusion implementation is that the runtime exposes no timestep, noise schedule, or loop over denoising steps. However, the first 1H pre-block internally synthesizes Gaussian-like values from a seed-like scalar and tile coordinates, and those values enter the first tensor-core computation. The best inference-level classification is therefore one-pass latent-conditioned generative rendering.
flowchart LR
subgraph raw[Engine and temporal resources]
color["Color / RGB"]
mv["MVec + scale"]
depth["Depth"]
prev["dlssnr_prev_output"]
control["ControlMask / UI / Backbuffer"]
end
submit["CG2R network-manager submission\nresource packing + CUBIN bindings"]
pre["1H pre-block / input adapter\nRGB texture + feature packing"]
rng["Internal Gaussian path\nseed-like field + tile coordinates\n3 generated FP16 lanes + 1.0"]
e1["1H / 32\ntiled Swin"]
e2["2H / 64\ntiled Swin"]
e3["4H / 128\ntiled Swin"]
e4["8H / 256\ntiled Swin"]
core["16H split-Swin compound core\nFFN + fused QKVAttn + projections/pool\nfinal head: 512 -> 1024"]
bridge["CCDecInputUpsample\n1024 -> 512\nmain + skip\noutput + reduce/overlap/conv-tap CBs"]
d4["8H decoder\nupsample + skip"]
d3["4H decoder\nupsample + skip"]
d2["2H decoder\nupsample + skip"]
d1["1H decoder\nupsample + skip"]
post["1H post-block\n2 inputs -> CudaSurface\nmask / simple-blend variants"]
output["Output\nplus history update"]
subgraph islands["Registered ViT islands; active placement unresolved"]
vit2d["ViT 2D compound block\nFFN + QKV + attention + projection"]
vit1d["ViT 1D compound block\nFFN + QKV + attention + projection"]
repack["2D <-> 1D repack"]
vit2d -.-> vit1d -.-> repack
end
color --> submit
mv --> submit
depth --> submit
prev --> submit
submit --> pre
rng -. generated inside pre-block .-> pre
pre --> e1 --> e2 --> e3 --> e4 --> core --> bridge
e4 -. skip .-> d4
e3 -. skip .-> d3
e2 -. skip .-> d2
e1 -. skip .-> d1
bridge --> d4 --> d3 --> d2 --> d1 --> post --> output
control --> post
output -. writes .-> prev
classDef observed fill:#d9f2e6,stroke:#24734d,color:#123b29;
classDef inferred fill:#fff1cc,stroke:#9a6b00,color:#4d3500;
class submit,pre,rng,e1,e2,e3,e4,core,bridge,d4,d3,d2,d1,post observed;
class raw,prev,output,vit2d,vit1d,repack inferred;
The solid path is a family-level skeleton, not a claim that each displayed
level maps to exactly one serialized blockN
. The weight resource contains many repeated records and sub-block tensors. The dashed skip edges are structurally supported by the decoder kernels and wrapper contracts, but the exact merge operation is not yet separated from the fused code.
The ViT/1D blocks are shown as a separate registered island because their internal inventories are clear while the active descriptor sequence does not yet establish whether they sit before the 16H core, between the core and decoder, or in a parallel conditioning branch.
| property | recovered value | confidence |
|---|---|---|
| target DLL | NVIDIA NGX rel_310_8 , nvngx_dlssnr.dll |
|
| high | ||
| SHA-256 | e16bcf15e16e13f527491cdf7845b2fe6521a738d8f7c9c721866a8496e1fc8e |
|
| high | ||
| learned-resource payload | WEIGHTS_HT/1033 , 147,695,410 bytes |
|
| high | ||
| resource records | exactly 153 parsed records across 71 logical block IDs (block0 –block70 ) |
|
| high | ||
| embedded CUDA | 15 fatbins and 231 public-inventory kernel names | high |
| GPU target | sm_120 / Blackwell only in this build |
|
| high | ||
| architecture | hierarchical tiled Swin + split-Swin/ViT/1D compound runtime + decoder routing | high for vocabulary; medium for exact topology |
| visible hierarchy | 1H/32 → 2H/64 → 4H/128 → 8H/256, with a 16H/512 central core | high for families; medium for placement/counts |
| central compound block | six specialized layer types, including fused QKV-attention and a final width expansion | high |
| decoder bridge | CCDecInputUpsample , 1024 → 512, main + skip |
|
| high | ||
| parameter count | approximately 148 million bytes of opaque learned-resource payload; exact independent scalar count remains unresolved | medium |
| temporal path | previous output bound into the network-manager submission and updated after evaluation | high structurally; medium for exact slot |
| internal random path | hash/coordinate-derived Gaussian-like FP16 lanes in the 1H pre-block | high for mechanics; medium for semantics |
| inference style | one forward submission; no visible timestep or denoising loop | high |
The public write-up's estimate of roughly 148 million FP8 parameters is
credible as a rounded model-size figure, but the static resource parser now
gives a more precise statement of what is actually measured. The
WEIGHTS_HT/1033
resource is 147,695,410 bytes and contains exactly 153 named records. Their opaque payloads total 147,683,778 bytes; the remaining 11,632 bytes are the resource header, record metadata, padding, and footer. The record tags are raw storage markers (including several non-printable-looking values), not yet decoded E4M3/E5M3 labels. Some records also contain 2-byte or small auxiliary payloads. Consequently, the payload total is not yet an exact count of independent learned scalars.
What is now exact at the resource level is the named-record and block inventory: 71 logical IDs, with a near-symmetric 4/5/8-layer core pattern. That inventory materially constrains the topology even though the active descriptor order and storage decoding remain separate questions.
The resource parser follows the record framing and validates the resource-size
field against the actual 147,695,410-byte blob. It recovers the public
write-up's 153-record count, including the small block70.layer0.blend_scale
record that is easy to miss if only blockN.layerM.layer
names are counted.
The resource contains 71 numeric block IDs (block0
through block70
). The most informative aggregate pattern is:
| block IDs | records per block | payload bytes per block | structural reading |
|---|---|---|---|
block0 –block4 |
|||
| 1 | 20–23 KiB | edge/adapter-size records | |
block5 –block8 |
|||
| 1 | 62–70 KiB | small transition records | |
block9 –block14 |
|||
| 1 | 197–230 KiB | medium transition records | |
block15 –block22 |
|||
| 1 | 689–820 KiB | wide single-layer records | |
block23 –block29 |
|||
| 4 | 1,968,192 | repeated 4-layer compound group | |
block30 |
|||
| 5 | 2,492,496 | 5-layer compound boundary group | |
block31 –block38 |
|||
| 5 | 12,587,154 | eight repeated central 5-layer groups | |
block39 |
|||
| 1 | 525,312 | central transition record | |
block40 –block47 |
|||
| 4 | 1,968,192 | repeated 4-layer compound group | |
block48 –block55 |
|||
| 1 | 689–821 KiB | wide single-layer records | |
block56 –block61 |
|||
| 1 | 197–230 KiB | medium transition records | |
block62 –block69 |
|||
| 1 | 21–70 KiB | small transition/adapter records | |
block70 |
|||
| 2 | 21,810 | 2-byte blend_scale plus one layer record |
The ranges above intentionally summarize similar payload sizes rather than assigning each ID to an encoder or decoder stage. The center's eight repeated five-record groups, flanked by repeated four-record groups, are strong structural corroboration for a large compound core surrounded by hierarchical transitions. They do not, by themselves, prove the runtime's active execution order or map one serialized block to one displayed Swin family.
The feature wrapper exposes DLSSNR.Color
, MVec
, and Depth
as core inputs,
with optional ControlMask
, UI
, UIAlpha
, Backbuffer
, and
BidirectionalDistortionField
resources. It also exposes controls including
Intensity
, LocalToneStrength
, LocalStructureStrength
,
SkinStructureStrength
, UseAutoMask
, UICorrection
, Enabled
, Reset
,
and DepthInverted
.
The runtime maintains separate dlssnr_original_color
and
dlssnr_network_output_scratch
surfaces in addition to dlssnr_prev_output
.
The evaluator creates the scratch/original-color resources and invokes
CG2RNetworkManager::Evaluate
(sub_180021bb0
). The manager reads the feature's previous-output field and sends it through the CUBIN resource-binding path. A later copy/update path writes the persistent history resource.
This is strong evidence for recurrent history feedback into the learned network submission rather than a history buffer used only by final blending. The exact history representation, reprojection, and insertion point remain open: the current evidence does not prove whether history is concatenated at the pre-block, injected deeper in the hierarchy, or transformed by a dedicated temporal kernel.
The CPU runtime registers 1H, 2H, 4H, and 8H fused Swin families. Representative CUDA modules use tensor-core MMA, explicit shared-memory staging, and variants for input/output views, chaining, waiting, tile synchronization, upsampling, and down/scale routing. The family names encode widths and head counts:
| family | visible width/head clue | representative role |
|---|---|---|
cc_tinlayout_fused_pre_block_swin_1h_32* |
||
| 1H / 32-wide | RGB-aware pre-block and first fused block | |
cc_tinlayout_fused_swin_1h_32* |
||
| 1H / 32-wide | regular fine-resolution local block | |
cc_tinlayout_fused_swin_2h_64_2* |
||
| 2H / 64-wide | hierarchical local block | |
cc_tinlayout_fused_swin_4h_128_4* |
||
| 4H / 128-wide | hierarchical local block | |
cc_tinlayout_fused_swin_8h_256_8* |
||
| 8H / 256-wide | hierarchical local block and decoder counterpart |
The pre-block's template includes FusedSwin2d1HConfig<32,...>
and
MpCubicSiluActivation
. The regular 1H body has 512 MMA operations and
repeated rsqrt
operations; the representative 2H, 4H, and 8H bodies have approximately 512, 448, and 576 MMA operations respectively. Non-FP8 shared footprints grow from roughly 9 KiB to 17 KiB to 34 KiB across those wider families, consistent with larger local tiles and more parallel channels.
The decoder-specific variants expose _upsample
names and the host runtime
has an explicit cc_tinlayout_upsample_skip_block
factory. This supports a U-Net-like hierarchy with encoder features retained for decoder-side routing, although the exact concat/add/gated merge is unresolved.
The host wrapper CCSplitSwin16HBlock
(sub_180040260
) recognizes six specialized layer types:
CCSplitSwin16HFfwd
CCSplitSwin16HFfwdProj
CCSplitSwin16HQKVAttn
CCSplitSwin16HProj
CCSplitSwin16HProjPool
CCSplitSwin16HFinalHead
The corresponding CUDA family contains separate FFN, FFN-projection, QKV,
projection, projection+pool, and final-head entries. The host type is
explicitly QKVAttn
, and there is no standalone split-Swin attention entry; the QKV family owns the fused middle operation. The base FFN entry contains 256 MMA operations, while the QKV entry contains 224 MMA operations plus warp-level reduction, normalization, and reciprocal work.
The final-head entry is a central feature transition, not the final image output head. Its template is:
Conv2d1x1Config<1024,512,...>
The first convolution template channel is the output width, as independently
confirmed by the ViT FFN expand/contract pair. Therefore this layer expands
512 → 1024. The later decoder bridge uses Conv2d1x1Config<512,1024,...>
and
contracts 1024 → 512. The host CCDecInputUpsample
constructor independently
checks 0x400 → 0x200
and identifies its specialization as “1024->512 (dec5)”.
The host-side layer contracts make the split-Swin interface more specific than
the kernel names alone. CCSplitSwin16HFfwd
is a one-input/one-output layer;
CCSplitSwin16HFfwdProj
takes (skip, src)
and produces one output;
CCSplitSwin16HQKVAttn
is one input/one output; CCSplitSwin16HProj
takes
(src, skip)
and produces one output; CCSplitSwin16HProjPool
takes two
inputs and produces two outputs; and CCSplitSwin16HFinalHead
is one input/one output. This is consistent with an internally routed compound block, rather than six independent top-level network stages.
The 2D ViT wrapper (sub_180041ca0
) recognizes five layer types:
CCVitFfnExpand
CCVitFfnContract
CCVitQKV
CCVitAttention
CCVitProjection
The 1D wrapper (sub_180043260
) recognizes corresponding 1D variants and
explicitly checks for five layer descriptors. The CUDA family adds
cc_vit_1d_repack_2d_to_1d
and cc_vit_1d_repack_1d_to_2d
entries.
Representative ViT QKV entries contain 192 MMA operations, consistent with three projections. Separate attention entries contain 128 MMA operations for QK-like work and weighted V accumulation. FFN expand/contract entries are separate; the 2D expand/contract templates are 1024 → 4096 and 4096 → 1024.
These blocks are high-confidence runtime capabilities and compound structures, but their active placement remains unresolved because the generic factory and constructor paths expose type dispatch rather than the serialized active network sequence.
The following contracts come from the host wrappers' validation strings and are more reliable than inferring interfaces from register allocation alone:
| layer family | inputs | outputs / side buffers | confidence |
|---|---|---|---|
| 1H pre-block | RGB texture input | one output; _ds exposes (pool, swin) |
|
| high | |||
| 1H/2H/4H/8H regular Swin | one tensor; upsample variant additionally takes skip |
||
one output; _ds exposes (pool, swin) |
|||
| high | |||
| split-Swin FFN / FFN-projection | one tensor / (skip, src) |
||
| one output | high | ||
| split-Swin QKV-attention | one tensor | one output | high |
| split-Swin projection | (src, skip) |
||
| one output | high | ||
| split-Swin projection+pool | two tensors | two outputs, likely pool plus transformed projection | high for arity; medium for identity |
| split-Swin final head | one tensor | one tensor, 512 → 1024 | high |
| 2D ViT FFN expand | one tensor | one output | high |
| 2D ViT FFN contract | (skip, src) |
||
dst plus reduce CB |
|||
| high | |||
| 2D ViT QKV | one tensor | qry , key , val plus reduce CB |
|
| high | |||
| 2D ViT attention | qry , key , val |
||
| one output | high | ||
| 2D ViT projection | (src, skip) |
||
dst plus reduce CB |
|||
| high | |||
| 1D ViT FFN/QKV/attention/projection | explicit token tensors and input/output CBs | explicit Q/K/V and destination/reduction CBs | high |
| decoder input upsample | (main, skip) |
||
{tensor, reduce_cb, overlap_cb, conv_tap} |
|||
| high | |||
| post-block | (main, enc0 skip) |
||
one CudaSurface output |
|||
| high |
The ViT QKV contract is especially informative: the runtime really does
materialize three logical projections (qry
, key
, val
) plus a reduction control buffer at the wrapper boundary, even though the GPU implementation fuses parts of their processing. The 1D wrapper likewise names explicit input/output CBs for the token path and exposes 2D↔1D repack kernels. This narrows the possible model interfaces substantially without claiming that these ViT islands are active in the shipping forward path.
The specialized CCDecInputUpsample
wrapper checks the channel transition
1024 → 512
and requires two inputs named (main, skip)
. It also validates a four-part output bundle:
{ tensor, reduce_cb, overlap_cb, conv_tap }
The matching normal decoder kernel is a no-activation
Conv2d1x1Config<512,1024,...>
implementation with 64 MMA operations, async
shared staging, and repeated vectorized output stores. The FP8 variant changes
the storage tile configuration but preserves the same learned 1x1 bridge
role. The PTX shows global reduction updates and a final integer release/store
path indexed by tile/block coordinates, which explains why the wrapper exposes
reduction and overlap control buffers in addition to the tensor result. The
exact parameter-slot identity of reduce_cb
, overlap_cb
, and conv_tap
is not fully proven because the synchronization variants add pointer arguments, but the four-output contract itself is explicit.
The post kernel family is likewise more than a generic image store. The base
CCTinlayoutFusedPostBlockSwin1HLayer
wrapper requires (main, enc0 skip)
and
one CudaSurface
output. Its kernel variants include simple_blend
,
control_mask
, and full_rect
forms, with corresponding FP8 variants. The PTX has a late texture-sampling/control path and a four-component surface write; the most defensible interpretation is three transformed output lanes plus a fourth alpha/control-like lane, followed by mode-dependent blending. The exact color-space transform and history-update equation remain open.
The CUDA templates name the activation MpCubicSiluActivation
. The emitted code does not call a conventional transcendental sigmoid. Instead, the representative low-precision path clamps the input and evaluates a cubic-like piecewise approximation:
t = clamp(x, -4, +4)
p = (-0.0559082) * abs(t) + 0.447266
a = t * p + 0.894531
y = x * a
The constants are rounded to half precision in the recovered path. This is a hardware-friendly SiLU-like nonlinearity that keeps the fused blocks tensor- core oriented.
The representative hierarchical Swin and split-QKV kernels contain FP32
reductions followed by rsqrt.approx.ftz.f32
, gain/bias-like multiplies, and reciprocal operations. This establishes an explicit variance- or energy-style normalization stage. The binary evidence available so far does not prove whether every instance subtracts a mean like LayerNorm or instead uses an RMS/L2-style normalization, so the exact normalization equation remains open.
The 2D and 1D ViT attention families are not simple unfused matrix-multiply
wrappers. A representative cc_vit_1d_attention
path performs QK-like MMA,
applies a half2 affine/clamp transform, constructs a fast positive
exponential-like surrogate using bit manipulation, clamps the denominator to
approximately 6.2e-5
, and accumulates V with reciprocal lanes. It is best
described as softmax-like attention; the recovered code does not expose a
canonical library softmax or ex2.approx
sequence.
The central split-Swin QKV family similarly fuses projection, reduction,
normalization, reciprocal, and middle-operation work. Because the host type is
CCSplitSwin16HQKVAttn
and there is no separate split-attention registration, the exact split attention/mixer equation cannot yet be isolated from its QKV kernel.
The embedded code targets Blackwell sm_120
and includes tensor-core FP8
conversion instructions such as UF2FP.SATFINITE.E4M3.F16
and
F2FP.SATFINITE.E4M3.F16
; both E4M3 and E5M3 strings are present. The _fp8
variants are the explicit low-precision weight path. Other suffixes encode execution variants rather than new neural layers:
| suffix/family | likely role |
|---|---|
_fp8 |
|
| FP8 weight/dequantization path | |
_chained , _wait , _tilesync |
|
| dependency and tile-scheduling variants | |
_inpview , _outview |
|
| input/output view-layout variants | |
_upsample , *_ds |
|
| hierarchy transition or decoder variants | |
repack_2d_to_1d , repack_1d_to_2d |
|
| ViT 2D/1D layout conversion | |
control_mask , simple_blend , full_rect |
|
| post-block control/output modes |
This naming matrix is one reason the public kernel count is much larger than the number of conceptual neural layers.
The strongest generative clue is inside the 1H pre-block rather than in a
caller-visible input surface. The raw PTX for the representative pre entry
contains a 264-byte parameter block, with a scalar at offset +200
and
dimension-like values at +208
. The scalar is multiplied by
-1640531527
, then combined with tile coordinates in a hash-like sequence.
The resulting code consumes four uniform-like values and executes the following recognizable pattern:
- transform uniform values with
lg2
,ln2
,-2
, andsqrt
; - use
sin
andcos
with a2π
-like constant, as in a Box–Muller-style Gaussian construction; - convert generated values to FP16;
- combine three generated lanes, a constant
1.0
, and four sampled texture lanes; - stage the resulting vector through shared memory before the first MMA.
The normal, FP8, and downsampled pre variants share this signature. No caller-provided noise surface, timestep, noise schedule, or loop over denoising steps was recovered. The safest interpretation is therefore an internal deterministic coordinate/seed-derived Gaussian-like conditioning path. Its exact semantic role—latent injection, stochastic feature augmentation, or a specialized preconditioning operation—remains unresolved.
The inference binary strongly supports a single-pass latent-conditioned generative renderer, but it does not contain the training metadata needed to name the objective. The current evidence ranks the interpretations as follows:
| interpretation | assessment from the binary |
|---|---|
| conventional multi-step diffusion sampler | unlikely: no timestep/schedule state or denoising loop is visible in the analyzed path |
| one-step distilled diffusion, consistency, flow/meanflow, or DMD-like model | strongly compatible with the one-pass latent-conditioned body and internal Gaussian-like path |
| GAN-style single-pass generator | also compatible with the observed inference graph |
| deterministic neural renderer with a random-looking internal conditioner | cannot be ruled out from implementation alone |
DLSS5 appears to combine several performance choices that are mutually reinforcing:
- Blackwell-only
sm_120
kernels use tensor-core MMA paths and explicit FP8 conversion variants. - Windowed/tiled Swin-style operators keep attention local and fuse nearby projection, normalization, activation, and layout work.
- The 16H compound core specializes its FFN, QKV-attention, projection, pool, and feature-transition stages instead of dispatching generic primitives.
- Async dependency variants (
_chained
,_wait
,_tilesync
) let the runtime match kernel scheduling to the tile graph. - A single analyzed forward path avoids the cost of a conventional denoising loop.
- Recurrent history feedback supplies temporal context without requiring a large sliding window of prior frames.
The current architecture is strong enough for a report, but several details would require live execution traces, more host-side state recovery, or weight interpretation:
- the serialized active layer sequence, repetition counts, and exact placement of the 2D/1D ViT islands;
- the meaning and lifetime of the
+200
pre-block scalar and the precise mapping from generated lanes to network channels; - the pre-block's complete RGB/motion/depth packing and the exact shape of its internal Gaussian-like injection;
- the exact mean-subtracting-vs-RMS normalization equation;
- the full split-Swin QKV-attention/mixer equation;
- whether decoder skip routing is add, concatenate, gated blend, or a fused
equivalent, and the exact identity of the decoder
reduce_cb
,overlap_cb
, andconv_tap
pointer slots; - the precise post-block color transform and output/history update semantics for each mask and blend mode;
- the active descriptor's exact block sequence and the placement/repetition of the ViT/1D islands. The builder selects a runtime descriptor and dispatches through factories, but the serialized descriptor contents are not exposed as a simple string list in this build;
- the training provenance: one-step diffusion distillation, consistency/flow matching, DMD-like regression, GAN, or another objective.
The following host-side procedures were the most useful anchors in Hopper:
| location | evidence recovered |
|---|---|
sub_18003c280 |
|
| family registration for pre/post 1H, hierarchical Swin, split-Swin, ViT, decoder bridge, and clear callback | |
sub_1800326c0 |
|
weight binding names: input_adapter_weight , weight1 , weight2 , ffn_cos_skip , qkv_weight , attn_scale , attn_bias , projection_weight , attn_cos_skip |
|
sub_18003cc80 |
|
| pre-block dispatch and the explicit “pre-block requires rgb input” guard | |
sub_18003fdc0 |
|
| single-layer dispatcher for 1H/2H/4H/8H Swin, fused pre/post blocks, and decoder input upsample | |
sub_180040260 |
|
six-layer CCSplitSwin16HBlock compound wrapper |
|
sub_180041ca0 |
|
five-layer 2D CCVitBlock wrapper |
|
sub_180043260 |
|
five-layer 1D CCVit1DBlock wrapper and repack expectations |
|
sub_180073bf0 |
|
| decoder bridge dimension checks and “1024->512 (dec5)” specialization | |
sub_18001f570 range |
|
| active-network builder: descriptor selection, runtime dimension logging, consolidated aligned weight-heap sizing, and factory-driven network construction | |
sub_180021bb0 |
|
CG2RNetworkManager::Evaluate , resource setup, previous-output handling, and CUBIN binding path |
|
tools/parse_dlss5_weight_map.py |
|
validated 153-record WEIGHTS_HT map, 71 logical block IDs, payload totals, and raw storage tags |
The CUDA-side evidence comes from the 15 embedded sm_120
fatbins, the representative PTX/SASS reductions, and the kernel-family naming matrix. Static analysis establishes capabilities and likely dataflow; it does not by itself prove the runtime's active descriptor sequence.
| aspect | DLSS 4.5 report | DLSS5 finding |
|---|---|---|
| core organization | compact local-transformer/U-Net-like body | heterogeneous hierarchical Swin + 16H split-Swin core + registered ViT islands |
| attention/mixing | local softmax attention followed by a softmax-free local mixer | fused local Swin families, split QKV-attention, and ViT softmax-like attention |
| activation | cubic SiLU approximation | same named MpCubicSiluActivation family is present |
| normalization | more fully characterized from the smaller graph | rsqrt/gain-style normalization recovered; exact mean-vs-RMS form remains open |
| temporal path | previous output/history participates in the graph | recurrent previous-output/history path is again visible, with no evidence of a large sliding window |
| output path | dynamic anisotropic Gaussian reconstruction filter was recovered | post-block mask/simple-blend/full-rect families are present, but a comparable output filter has not been recovered |
| latent/random path | no analogous internal Gaussian signature reported | pre-block contains an internal Gaussian-like, coordinate/seed-derived path |
| inference interpretation | deterministic neural upscaler with specialized local mixing | single-pass latent-conditioned generative neural renderer |
| training claim | architecture analysis did not require a training-objective claim | one-step distilled diffusion is plausible, but not distinguishable from GAN/flow/consistency alternatives from the DLL alone |
The comparison is intentionally asymmetric: DLSS4.5's smaller model was amenable to a more complete graph reconstruction, while the DLSS5 artifact exposes a larger, more heterogeneous runtime with several compound islands whose active ordering is not serialized in the public evidence we have.
DLSS5 is best described as a recurrent, single-pass, latent-conditioned generative neural renderer. Its body combines a 1H input adapter, hierarchical 1H/2H/4H/8H tiled Swin stages, a 16H split-Swin compound core with a 512 → 1024 feature expansion, a 1024 → 512 decoder bridge, decoder-side upsample/ skip blocks, and a 1H mask/blend post path. ViT and 1D transformer blocks are also registered as specialized compound islands, although their active placement is not yet proven.
The internal coordinate/seed-derived Gaussian-like path and the absence of a runtime timestep or denoising loop make one-step generative inference a strong architectural interpretation. They do not, however, identify whether NVIDIA trained it with meanflow, consistency distillation, DMD, GAN losses, or some other teacher/student objective. That distinction belongs in the “likely training neighborhood,” not in the list of facts directly recovered from the binary.