{"slug": "what-is-the-dlss5-model-architecture", "title": "What is the DLSS5 model architecture?", "summary": "A developer reverse-engineered NVIDIA's DLSS5 binary to infer its model architecture, revealing a heterogeneous, hierarchical, recurrent, single-pass neural renderer with tiled Swin transformer families and a compound 16H split-Swin core. The analysis, published in a gist, indicates the runtime exposes no timestep or denoising loop, classifying the model as one-pass latent-conditioned generative rendering.", "body_md": "Let's attempt to figure out the model architecture from DLSS5's binary! See [https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee](https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee) for context.\n\nWarning\n\nThe remainder of this gist was authored by Codex 5.6\n\nThe strongest current model is a **heterogeneous, hierarchical, recurrent,\nsingle-pass neural renderer**. Its visible body combines tiled 1H/2H/4H/8H\nSwin families, a compound 16H split-Swin core, 2D and 1D ViT compound blocks,\nspecialized 1x1 width transitions, and a decoder with upsample/skip routing.\n\nThe most important difference from a conventional diffusion implementation is that the runtime exposes no timestep, noise schedule, or loop over denoising steps. However, the first 1H pre-block internally synthesizes Gaussian-like values from a seed-like scalar and tile coordinates, and those values enter the first tensor-core computation. The best inference-level classification is therefore one-pass latent-conditioned generative rendering.\n\n```\nflowchart LR\n  subgraph raw[Engine and temporal resources]\n    color[\"Color / RGB\"]\n    mv[\"MVec + scale\"]\n    depth[\"Depth\"]\n    prev[\"dlssnr_prev_output\"]\n    control[\"ControlMask / UI / Backbuffer\"]\n  end\n\n  submit[\"CG2R network-manager submission\\nresource packing + CUBIN bindings\"]\n  pre[\"1H pre-block / input adapter\\nRGB texture + feature packing\"]\n  rng[\"Internal Gaussian path\\nseed-like field + tile coordinates\\n3 generated FP16 lanes + 1.0\"]\n  e1[\"1H / 32\\ntiled Swin\"]\n  e2[\"2H / 64\\ntiled Swin\"]\n  e3[\"4H / 128\\ntiled Swin\"]\n  e4[\"8H / 256\\ntiled Swin\"]\n  core[\"16H split-Swin compound core\\nFFN + fused QKVAttn + projections/pool\\nfinal head: 512 -> 1024\"]\n  bridge[\"CCDecInputUpsample\\n1024 -> 512\\nmain + skip\\noutput + reduce/overlap/conv-tap CBs\"]\n  d4[\"8H decoder\\nupsample + skip\"]\n  d3[\"4H decoder\\nupsample + skip\"]\n  d2[\"2H decoder\\nupsample + skip\"]\n  d1[\"1H decoder\\nupsample + skip\"]\n  post[\"1H post-block\\n2 inputs -> CudaSurface\\nmask / simple-blend variants\"]\n  output[\"Output\\nplus history update\"]\n\n  subgraph islands[\"Registered ViT islands; active placement unresolved\"]\n    vit2d[\"ViT 2D compound block\\nFFN + QKV + attention + projection\"]\n    vit1d[\"ViT 1D compound block\\nFFN + QKV + attention + projection\"]\n    repack[\"2D <-> 1D repack\"]\n    vit2d -.-> vit1d -.-> repack\n  end\n\n  color --> submit\n  mv --> submit\n  depth --> submit\n  prev --> submit\n  submit --> pre\n  rng -. generated inside pre-block .-> pre\n  pre --> e1 --> e2 --> e3 --> e4 --> core --> bridge\n  e4 -. skip .-> d4\n  e3 -. skip .-> d3\n  e2 -. skip .-> d2\n  e1 -. skip .-> d1\n  bridge --> d4 --> d3 --> d2 --> d1 --> post --> output\n  control --> post\n  output -. writes .-> prev\n\n  classDef observed fill:#d9f2e6,stroke:#24734d,color:#123b29;\n  classDef inferred fill:#fff1cc,stroke:#9a6b00,color:#4d3500;\n  class submit,pre,rng,e1,e2,e3,e4,core,bridge,d4,d3,d2,d1,post observed;\n  class raw,prev,output,vit2d,vit1d,repack inferred;\n```\n\nThe solid path is a family-level skeleton, not a claim that each displayed\nlevel maps to exactly one serialized `blockN`\n\n. The weight resource contains\nmany repeated records and sub-block tensors. The dashed skip edges are\nstructurally supported by the decoder kernels and wrapper contracts, but the\nexact merge operation is not yet separated from the fused code.\n\nThe ViT/1D blocks are shown as a separate registered island because their internal inventories are clear while the active descriptor sequence does not yet establish whether they sit before the 16H core, between the core and decoder, or in a parallel conditioning branch.\n\n| property | recovered value | confidence |\n|---|---|---|\n| target DLL | NVIDIA NGX `rel_310_8` , `nvngx_dlssnr.dll` |\nhigh |\n| SHA-256 | `e16bcf15e16e13f527491cdf7845b2fe6521a738d8f7c9c721866a8496e1fc8e` |\nhigh |\n| learned-resource payload | `WEIGHTS_HT/1033` , 147,695,410 bytes |\nhigh |\n| resource records | exactly 153 parsed records across 71 logical block IDs (`block0` –`block70` ) |\nhigh |\n| embedded CUDA | 15 fatbins and 231 public-inventory kernel names | high |\n| GPU target | `sm_120` / Blackwell only in this build |\nhigh |\n| architecture | hierarchical tiled Swin + split-Swin/ViT/1D compound runtime + decoder routing | high for vocabulary; medium for exact topology |\n| visible hierarchy | 1H/32 → 2H/64 → 4H/128 → 8H/256, with a 16H/512 central core | high for families; medium for placement/counts |\n| central compound block | six specialized layer types, including fused QKV-attention and a final width expansion | high |\n| decoder bridge | `CCDecInputUpsample` , 1024 → 512, main + skip |\nhigh |\n| parameter count | approximately 148 million bytes of opaque learned-resource payload; exact independent scalar count remains unresolved | medium |\n| temporal path | previous output bound into the network-manager submission and updated after evaluation | high structurally; medium for exact slot |\n| internal random path | hash/coordinate-derived Gaussian-like FP16 lanes in the 1H pre-block | high for mechanics; medium for semantics |\n| inference style | one forward submission; no visible timestep or denoising loop | high |\n\nThe public write-up's estimate of roughly 148 million FP8 parameters is\ncredible as a rounded model-size figure, but the static resource parser now\ngives a more precise statement of what is actually measured. The\n`WEIGHTS_HT/1033`\n\nresource is 147,695,410 bytes and contains exactly 153 named\nrecords. Their opaque payloads total 147,683,778 bytes; the remaining 11,632\nbytes are the resource header, record metadata, padding, and footer. The\nrecord tags are raw storage markers (including several non-printable-looking\nvalues), not yet decoded E4M3/E5M3 labels. Some records also contain 2-byte or\nsmall auxiliary payloads. Consequently, the payload total is not yet an exact\ncount of independent learned scalars.\n\nWhat *is* now exact at the resource level is the named-record and block\ninventory: 71 logical IDs, with a near-symmetric 4/5/8-layer core pattern.\nThat inventory materially constrains the topology even though the active\ndescriptor order and storage decoding remain separate questions.\n\nThe resource parser follows the record framing and validates the resource-size\nfield against the actual 147,695,410-byte blob. It recovers the public\nwrite-up's 153-record count, including the small `block70.layer0.blend_scale`\n\nrecord that is easy to miss if only `blockN.layerM.layer`\n\nnames are counted.\nThe resource contains 71 numeric block IDs (`block0`\n\nthrough `block70`\n\n). The\nmost informative aggregate pattern is:\n\n| block IDs | records per block | payload bytes per block | structural reading |\n|---|---|---|---|\n`block0` –`block4` |\n1 | 20–23 KiB | edge/adapter-size records |\n`block5` –`block8` |\n1 | 62–70 KiB | small transition records |\n`block9` –`block14` |\n1 | 197–230 KiB | medium transition records |\n`block15` –`block22` |\n1 | 689–820 KiB | wide single-layer records |\n`block23` –`block29` |\n4 | 1,968,192 | repeated 4-layer compound group |\n`block30` |\n5 | 2,492,496 | 5-layer compound boundary group |\n`block31` –`block38` |\n5 | 12,587,154 | eight repeated central 5-layer groups |\n`block39` |\n1 | 525,312 | central transition record |\n`block40` –`block47` |\n4 | 1,968,192 | repeated 4-layer compound group |\n`block48` –`block55` |\n1 | 689–821 KiB | wide single-layer records |\n`block56` –`block61` |\n1 | 197–230 KiB | medium transition records |\n`block62` –`block69` |\n1 | 21–70 KiB | small transition/adapter records |\n`block70` |\n2 | 21,810 | 2-byte `blend_scale` plus one layer record |\n\nThe ranges above intentionally summarize similar payload sizes rather than assigning each ID to an encoder or decoder stage. The center's eight repeated five-record groups, flanked by repeated four-record groups, are strong structural corroboration for a large compound core surrounded by hierarchical transitions. They do not, by themselves, prove the runtime's active execution order or map one serialized block to one displayed Swin family.\n\nThe feature wrapper exposes `DLSSNR.Color`\n\n, `MVec`\n\n, and `Depth`\n\nas core inputs,\nwith optional `ControlMask`\n\n, `UI`\n\n, `UIAlpha`\n\n, `Backbuffer`\n\n, and\n`BidirectionalDistortionField`\n\nresources. It also exposes controls including\n`Intensity`\n\n, `LocalToneStrength`\n\n, `LocalStructureStrength`\n\n,\n`SkinStructureStrength`\n\n, `UseAutoMask`\n\n, `UICorrection`\n\n, `Enabled`\n\n, `Reset`\n\n,\nand `DepthInverted`\n\n.\n\nThe runtime maintains separate `dlssnr_original_color`\n\nand\n`dlssnr_network_output_scratch`\n\nsurfaces in addition to `dlssnr_prev_output`\n\n.\nThe evaluator creates the scratch/original-color resources and invokes\n`CG2RNetworkManager::Evaluate`\n\n(`sub_180021bb0`\n\n). The manager reads the\nfeature's previous-output field and sends it through the CUBIN resource-binding\npath. A later copy/update path writes the persistent history resource.\n\nThis is strong evidence for recurrent history feedback into the learned network submission rather than a history buffer used only by final blending. The exact history representation, reprojection, and insertion point remain open: the current evidence does not prove whether history is concatenated at the pre-block, injected deeper in the hierarchy, or transformed by a dedicated temporal kernel.\n\nThe CPU runtime registers 1H, 2H, 4H, and 8H fused Swin families. Representative CUDA modules use tensor-core MMA, explicit shared-memory staging, and variants for input/output views, chaining, waiting, tile synchronization, upsampling, and down/scale routing. The family names encode widths and head counts:\n\n| family | visible width/head clue | representative role |\n|---|---|---|\n`cc_tinlayout_fused_pre_block_swin_1h_32*` |\n1H / 32-wide | RGB-aware pre-block and first fused block |\n`cc_tinlayout_fused_swin_1h_32*` |\n1H / 32-wide | regular fine-resolution local block |\n`cc_tinlayout_fused_swin_2h_64_2*` |\n2H / 64-wide | hierarchical local block |\n`cc_tinlayout_fused_swin_4h_128_4*` |\n4H / 128-wide | hierarchical local block |\n`cc_tinlayout_fused_swin_8h_256_8*` |\n8H / 256-wide | hierarchical local block and decoder counterpart |\n\nThe pre-block's template includes `FusedSwin2d1HConfig<32,...>`\n\nand\n`MpCubicSiluActivation`\n\n. The regular 1H body has 512 MMA operations and\nrepeated `rsqrt`\n\noperations; the representative 2H, 4H, and 8H bodies have\napproximately 512, 448, and 576 MMA operations respectively. Non-FP8 shared\nfootprints grow from roughly 9 KiB to 17 KiB to 34 KiB across those wider\nfamilies, consistent with larger local tiles and more parallel channels.\n\nThe decoder-specific variants expose `_upsample`\n\nnames and the host runtime\nhas an explicit `cc_tinlayout_upsample_skip_block`\n\nfactory. This supports a\nU-Net-like hierarchy with encoder features retained for decoder-side routing,\nalthough the exact concat/add/gated merge is unresolved.\n\nThe host wrapper `CCSplitSwin16HBlock`\n\n(`sub_180040260`\n\n) recognizes six\nspecialized layer types:\n\n```\nCCSplitSwin16HFfwd\nCCSplitSwin16HFfwdProj\nCCSplitSwin16HQKVAttn\nCCSplitSwin16HProj\nCCSplitSwin16HProjPool\nCCSplitSwin16HFinalHead\n```\n\nThe corresponding CUDA family contains separate FFN, FFN-projection, QKV,\nprojection, projection+pool, and final-head entries. The host type is\nexplicitly `QKVAttn`\n\n, and there is no standalone split-Swin attention entry;\nthe QKV family owns the fused middle operation. The base FFN entry contains\n256 MMA operations, while the QKV entry contains 224 MMA operations plus\nwarp-level reduction, normalization, and reciprocal work.\n\nThe final-head entry is a central feature transition, not the final image output head. Its template is:\n\n```\nConv2d1x1Config<1024,512,...>\n```\n\nThe first convolution template channel is the output width, as independently\nconfirmed by the ViT FFN expand/contract pair. Therefore this layer expands\n512 → 1024. The later decoder bridge uses `Conv2d1x1Config<512,1024,...>`\n\nand\ncontracts 1024 → 512. The host `CCDecInputUpsample`\n\nconstructor independently\nchecks `0x400 → 0x200`\n\nand identifies its specialization as “1024->512\n(dec5)”.\n\nThe host-side layer contracts make the split-Swin interface more specific than\nthe kernel names alone. `CCSplitSwin16HFfwd`\n\nis a one-input/one-output layer;\n`CCSplitSwin16HFfwdProj`\n\ntakes `(skip, src)`\n\nand produces one output;\n`CCSplitSwin16HQKVAttn`\n\nis one input/one output; `CCSplitSwin16HProj`\n\ntakes\n`(src, skip)`\n\nand produces one output; `CCSplitSwin16HProjPool`\n\ntakes two\ninputs and produces two outputs; and `CCSplitSwin16HFinalHead`\n\nis one\ninput/one output. This is consistent with an internally routed compound\nblock, rather than six independent top-level network stages.\n\nThe 2D ViT wrapper (`sub_180041ca0`\n\n) recognizes five layer types:\n\n```\nCCVitFfnExpand\nCCVitFfnContract\nCCVitQKV\nCCVitAttention\nCCVitProjection\n```\n\nThe 1D wrapper (`sub_180043260`\n\n) recognizes corresponding 1D variants and\nexplicitly checks for five layer descriptors. The CUDA family adds\n`cc_vit_1d_repack_2d_to_1d`\n\nand `cc_vit_1d_repack_1d_to_2d`\n\nentries.\n\nRepresentative ViT QKV entries contain 192 MMA operations, consistent with three projections. Separate attention entries contain 128 MMA operations for QK-like work and weighted V accumulation. FFN expand/contract entries are separate; the 2D expand/contract templates are 1024 → 4096 and 4096 → 1024.\n\nThese blocks are high-confidence runtime capabilities and compound structures, but their active placement remains unresolved because the generic factory and constructor paths expose type dispatch rather than the serialized active network sequence.\n\nThe following contracts come from the host wrappers' validation strings and are more reliable than inferring interfaces from register allocation alone:\n\n| layer family | inputs | outputs / side buffers | confidence |\n|---|---|---|---|\n| 1H pre-block | RGB texture input | one output; `_ds` exposes `(pool, swin)` |\nhigh |\n| 1H/2H/4H/8H regular Swin | one tensor; upsample variant additionally takes `skip` |\none output; `_ds` exposes `(pool, swin)` |\nhigh |\n| split-Swin FFN / FFN-projection | one tensor / `(skip, src)` |\none output | high |\n| split-Swin QKV-attention | one tensor | one output | high |\n| split-Swin projection | `(src, skip)` |\none output | high |\n| split-Swin projection+pool | two tensors | two outputs, likely pool plus transformed projection | high for arity; medium for identity |\n| split-Swin final head | one tensor | one tensor, 512 → 1024 | high |\n| 2D ViT FFN expand | one tensor | one output | high |\n| 2D ViT FFN contract | `(skip, src)` |\n`dst` plus reduce CB |\nhigh |\n| 2D ViT QKV | one tensor | `qry` , `key` , `val` plus reduce CB |\nhigh |\n| 2D ViT attention | `qry` , `key` , `val` |\none output | high |\n| 2D ViT projection | `(src, skip)` |\n`dst` plus reduce CB |\nhigh |\n| 1D ViT FFN/QKV/attention/projection | explicit token tensors and input/output CBs | explicit Q/K/V and destination/reduction CBs | high |\n| decoder input upsample | `(main, skip)` |\n`{tensor, reduce_cb, overlap_cb, conv_tap}` |\nhigh |\n| post-block | `(main, enc0 skip)` |\none `CudaSurface` output |\nhigh |\n\nThe ViT QKV contract is especially informative: the runtime really does\nmaterialize three logical projections (`qry`\n\n, `key`\n\n, `val`\n\n) plus a reduction\ncontrol buffer at the wrapper boundary, even though the GPU implementation\nfuses parts of their processing. The 1D wrapper likewise names explicit\ninput/output CBs for the token path and exposes 2D↔1D repack kernels. This\nnarrows the possible model interfaces substantially without claiming that\nthese ViT islands are active in the shipping forward path.\n\nThe specialized `CCDecInputUpsample`\n\nwrapper checks the channel transition\n`1024 → 512`\n\nand requires two inputs named `(main, skip)`\n\n. It also validates a\nfour-part output bundle:\n\n```\n{ tensor, reduce_cb, overlap_cb, conv_tap }\n```\n\nThe matching normal decoder kernel is a no-activation\n`Conv2d1x1Config<512,1024,...>`\n\nimplementation with 64 MMA operations, async\nshared staging, and repeated vectorized output stores. The FP8 variant changes\nthe storage tile configuration but preserves the same learned 1x1 bridge\nrole. The PTX shows global reduction updates and a final integer release/store\npath indexed by tile/block coordinates, which explains why the wrapper exposes\nreduction and overlap control buffers in addition to the tensor result. The\nexact parameter-slot identity of `reduce_cb`\n\n, `overlap_cb`\n\n, and `conv_tap`\n\nis\nnot fully proven because the synchronization variants add pointer arguments,\nbut the four-output contract itself is explicit.\n\nThe post kernel family is likewise more than a generic image store. The base\n`CCTinlayoutFusedPostBlockSwin1HLayer`\n\nwrapper requires `(main, enc0 skip)`\n\nand\none `CudaSurface`\n\noutput. Its kernel variants include `simple_blend`\n\n,\n`control_mask`\n\n, and `full_rect`\n\nforms, with corresponding FP8 variants. The\nPTX has a late texture-sampling/control path and a four-component surface\nwrite; the most defensible interpretation is three transformed output lanes\nplus a fourth alpha/control-like lane, followed by mode-dependent blending.\nThe exact color-space transform and history-update equation remain open.\n\nThe CUDA templates name the activation `MpCubicSiluActivation`\n\n. The emitted\ncode does not call a conventional transcendental sigmoid. Instead, the\nrepresentative low-precision path clamps the input and evaluates a cubic-like\npiecewise approximation:\n\n```\nt = clamp(x, -4, +4)\np = (-0.0559082) * abs(t) + 0.447266\na = t * p + 0.894531\ny = x * a\n```\n\nThe constants are rounded to half precision in the recovered path. This is a hardware-friendly SiLU-like nonlinearity that keeps the fused blocks tensor- core oriented.\n\nThe representative hierarchical Swin and split-QKV kernels contain FP32\nreductions followed by `rsqrt.approx.ftz.f32`\n\n, gain/bias-like multiplies, and\nreciprocal operations. This establishes an explicit variance- or energy-style\nnormalization stage. The binary evidence available so far does not prove\nwhether every instance subtracts a mean like LayerNorm or instead uses an\nRMS/L2-style normalization, so the exact normalization equation remains open.\n\nThe 2D and 1D ViT attention families are not simple unfused matrix-multiply\nwrappers. A representative `cc_vit_1d_attention`\n\npath performs QK-like MMA,\napplies a half2 affine/clamp transform, constructs a fast positive\nexponential-like surrogate using bit manipulation, clamps the denominator to\napproximately `6.2e-5`\n\n, and accumulates V with reciprocal lanes. It is best\ndescribed as softmax-like attention; the recovered code does not expose a\ncanonical library softmax or `ex2.approx`\n\nsequence.\n\nThe central split-Swin QKV family similarly fuses projection, reduction,\nnormalization, reciprocal, and middle-operation work. Because the host type is\n`CCSplitSwin16HQKVAttn`\n\nand there is no separate split-attention registration,\nthe exact split attention/mixer equation cannot yet be isolated from its QKV\nkernel.\n\nThe embedded code targets Blackwell `sm_120`\n\nand includes tensor-core FP8\nconversion instructions such as `UF2FP.SATFINITE.E4M3.F16`\n\nand\n`F2FP.SATFINITE.E4M3.F16`\n\n; both E4M3 and E5M3 strings are present. The `_fp8`\n\nvariants are the explicit low-precision weight path. Other suffixes encode\nexecution variants rather than new neural layers:\n\n| suffix/family | likely role |\n|---|---|\n`_fp8` |\nFP8 weight/dequantization path |\n`_chained` , `_wait` , `_tilesync` |\ndependency and tile-scheduling variants |\n`_inpview` , `_outview` |\ninput/output view-layout variants |\n`_upsample` , `*_ds` |\nhierarchy transition or decoder variants |\n`repack_2d_to_1d` , `repack_1d_to_2d` |\nViT 2D/1D layout conversion |\n`control_mask` , `simple_blend` , `full_rect` |\npost-block control/output modes |\n\nThis naming matrix is one reason the public kernel count is much larger than the number of conceptual neural layers.\n\nThe strongest generative clue is inside the 1H pre-block rather than in a\ncaller-visible input surface. The raw PTX for the representative pre entry\ncontains a 264-byte parameter block, with a scalar at offset `+200`\n\nand\ndimension-like values at `+208`\n\n. The scalar is multiplied by\n`-1640531527`\n\n, then combined with tile coordinates in a hash-like sequence.\n\nThe resulting code consumes four uniform-like values and executes the following recognizable pattern:\n\n- transform uniform values with\n`lg2`\n\n,`ln2`\n\n,`-2`\n\n, and`sqrt`\n\n; - use\n`sin`\n\nand`cos`\n\nwith a`2π`\n\n-like constant, as in a Box–Muller-style Gaussian construction; - convert generated values to FP16;\n- combine three generated lanes, a constant\n`1.0`\n\n, and four sampled texture lanes; - stage the resulting vector through shared memory before the first MMA.\n\nThe normal, FP8, and downsampled pre variants share this signature. No caller-provided noise surface, timestep, noise schedule, or loop over denoising steps was recovered. The safest interpretation is therefore an internal deterministic coordinate/seed-derived Gaussian-like conditioning path. Its exact semantic role—latent injection, stochastic feature augmentation, or a specialized preconditioning operation—remains unresolved.\n\nThe inference binary strongly supports a single-pass latent-conditioned generative renderer, but it does not contain the training metadata needed to name the objective. The current evidence ranks the interpretations as follows:\n\n| interpretation | assessment from the binary |\n|---|---|\n| conventional multi-step diffusion sampler | unlikely: no timestep/schedule state or denoising loop is visible in the analyzed path |\n| one-step distilled diffusion, consistency, flow/meanflow, or DMD-like model | strongly compatible with the one-pass latent-conditioned body and internal Gaussian-like path |\n| GAN-style single-pass generator | also compatible with the observed inference graph |\n| deterministic neural renderer with a random-looking internal conditioner | cannot be ruled out from implementation alone |\n\nDLSS5 appears to combine several performance choices that are mutually reinforcing:\n\n- Blackwell-only\n`sm_120`\n\nkernels use tensor-core MMA paths and explicit FP8 conversion variants. - Windowed/tiled Swin-style operators keep attention local and fuse nearby projection, normalization, activation, and layout work.\n- The 16H compound core specializes its FFN, QKV-attention, projection, pool, and feature-transition stages instead of dispatching generic primitives.\n- Async dependency variants (\n`_chained`\n\n,`_wait`\n\n,`_tilesync`\n\n) let the runtime match kernel scheduling to the tile graph. - A single analyzed forward path avoids the cost of a conventional denoising loop.\n- Recurrent history feedback supplies temporal context without requiring a large sliding window of prior frames.\n\nThe current architecture is strong enough for a report, but several details would require live execution traces, more host-side state recovery, or weight interpretation:\n\n- the serialized active layer sequence, repetition counts, and exact placement of the 2D/1D ViT islands;\n- the meaning and lifetime of the\n`+200`\n\npre-block scalar and the precise mapping from generated lanes to network channels; - the pre-block's complete RGB/motion/depth packing and the exact shape of its internal Gaussian-like injection;\n- the exact mean-subtracting-vs-RMS normalization equation;\n- the full split-Swin QKV-attention/mixer equation;\n- whether decoder skip routing is add, concatenate, gated blend, or a fused\nequivalent, and the exact identity of the decoder\n`reduce_cb`\n\n,`overlap_cb`\n\n, and`conv_tap`\n\npointer slots; - the precise post-block color transform and output/history update semantics for each mask and blend mode;\n- the active descriptor's exact block sequence and the placement/repetition of the ViT/1D islands. The builder selects a runtime descriptor and dispatches through factories, but the serialized descriptor contents are not exposed as a simple string list in this build;\n- the training provenance: one-step diffusion distillation, consistency/flow matching, DMD-like regression, GAN, or another objective.\n\nThe following host-side procedures were the most useful anchors in Hopper:\n\n| location | evidence recovered |\n|---|---|\n`sub_18003c280` |\nfamily registration for pre/post 1H, hierarchical Swin, split-Swin, ViT, decoder bridge, and clear callback |\n`sub_1800326c0` |\nweight binding names: `input_adapter_weight` , `weight1` , `weight2` , `ffn_cos_skip` , `qkv_weight` , `attn_scale` , `attn_bias` , `projection_weight` , `attn_cos_skip` |\n`sub_18003cc80` |\npre-block dispatch and the explicit “pre-block requires rgb input” guard |\n`sub_18003fdc0` |\nsingle-layer dispatcher for 1H/2H/4H/8H Swin, fused pre/post blocks, and decoder input upsample |\n`sub_180040260` |\nsix-layer `CCSplitSwin16HBlock` compound wrapper |\n`sub_180041ca0` |\nfive-layer 2D `CCVitBlock` wrapper |\n`sub_180043260` |\nfive-layer 1D `CCVit1DBlock` wrapper and repack expectations |\n`sub_180073bf0` |\ndecoder bridge dimension checks and “1024->512 (dec5)” specialization |\n`sub_18001f570` range |\nactive-network builder: descriptor selection, runtime dimension logging, consolidated aligned weight-heap sizing, and factory-driven network construction |\n`sub_180021bb0` |\n`CG2RNetworkManager::Evaluate` , resource setup, previous-output handling, and CUBIN binding path |\n`tools/parse_dlss5_weight_map.py` |\nvalidated 153-record `WEIGHTS_HT` map, 71 logical block IDs, payload totals, and raw storage tags |\n\nThe CUDA-side evidence comes from the 15 embedded `sm_120`\n\nfatbins, the\nrepresentative PTX/SASS reductions, and the kernel-family naming matrix.\nStatic analysis establishes capabilities and likely dataflow; it does not by\nitself prove the runtime's active descriptor sequence.\n\n| aspect | DLSS 4.5 report | DLSS5 finding |\n|---|---|---|\n| core organization | compact local-transformer/U-Net-like body | heterogeneous hierarchical Swin + 16H split-Swin core + registered ViT islands |\n| attention/mixing | local softmax attention followed by a softmax-free local mixer | fused local Swin families, split QKV-attention, and ViT softmax-like attention |\n| activation | cubic SiLU approximation | same named `MpCubicSiluActivation` family is present |\n| normalization | more fully characterized from the smaller graph | rsqrt/gain-style normalization recovered; exact mean-vs-RMS form remains open |\n| temporal path | previous output/history participates in the graph | recurrent previous-output/history path is again visible, with no evidence of a large sliding window |\n| output path | dynamic anisotropic Gaussian reconstruction filter was recovered | post-block mask/simple-blend/full-rect families are present, but a comparable output filter has not been recovered |\n| latent/random path | no analogous internal Gaussian signature reported | pre-block contains an internal Gaussian-like, coordinate/seed-derived path |\n| inference interpretation | deterministic neural upscaler with specialized local mixing | single-pass latent-conditioned generative neural renderer |\n| training claim | architecture analysis did not require a training-objective claim | one-step distilled diffusion is plausible, but not distinguishable from GAN/flow/consistency alternatives from the DLL alone |\n\nThe comparison is intentionally asymmetric: DLSS4.5's smaller model was amenable to a more complete graph reconstruction, while the DLSS5 artifact exposes a larger, more heterogeneous runtime with several compound islands whose active ordering is not serialized in the public evidence we have.\n\nDLSS5 is best described as a recurrent, single-pass, latent-conditioned generative neural renderer. Its body combines a 1H input adapter, hierarchical 1H/2H/4H/8H tiled Swin stages, a 16H split-Swin compound core with a 512 → 1024 feature expansion, a 1024 → 512 decoder bridge, decoder-side upsample/ skip blocks, and a 1H mask/blend post path. ViT and 1D transformer blocks are also registered as specialized compound islands, although their active placement is not yet proven.\n\nThe internal coordinate/seed-derived Gaussian-like path and the absence of a runtime timestep or denoising loop make one-step generative inference a strong architectural interpretation. They do not, however, identify whether NVIDIA trained it with meanflow, consistency distillation, DMD, GAN losses, or some other teacher/student objective. That distinction belongs in the “likely training neighborhood,” not in the list of facts directly recovered from the binary.", "url": "https://wpnews.pro/news/what-is-the-dlss5-model-architecture", "canonical_source": "https://gist.github.com/madebyollin/55c703a34bf90962844edcd68d04e32e", "published_at": "2026-08-30 15:12:05+00:00", "updated_at": "2026-09-02 06:51:55.720466+00:00", "lang": "en", "topics": ["machine-learning", "computer-vision", "neural-networks", "ai-research", "ai-infrastructure"], "entities": ["NVIDIA", "DLSS5", "Codex 5.6", "Swin", "ViT"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-is-the-dlss5-model-architecture", "markdown": "https://wpnews.pro/news/what-is-the-dlss5-model-architecture.md", "text": "https://wpnews.pro/news/what-is-the-dlss5-model-architecture.txt", "jsonld": "https://wpnews.pro/news/what-is-the-dlss5-model-architecture.jsonld"}}