cd /news/machine-learning/what-is-the-dlss5-model-architecture · home › topics › machine-learning › article
[ARTICLE · art-118659] src=gist.github.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

What is the DLSS5 model architecture?

A developer reverse-engineered NVIDIA's DLSS5 binary to infer its model architecture, revealing a heterogeneous, hierarchical, recurrent, single-pass neural renderer with tiled Swin transformer families and a compound 16H split-Swin core. The analysis, published in a gist, indicates the runtime exposes no timestep or denoising loop, classifying the model as one-pass latent-conditioned generative rendering.

read20 min views49 publishedAug 30, 2026

Let's attempt to figure out the model architecture from DLSS5's binary! See https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee for context.

Warning

The remainder of this gist was authored by Codex 5.6

The strongest current model is a heterogeneous, hierarchical, recurrent, single-pass neural renderer. Its visible body combines tiled 1H/2H/4H/8H Swin families, a compound 16H split-Swin core, 2D and 1D ViT compound blocks, specialized 1x1 width transitions, and a decoder with upsample/skip routing.

The most important difference from a conventional diffusion implementation is that the runtime exposes no timestep, noise schedule, or loop over denoising steps. However, the first 1H pre-block internally synthesizes Gaussian-like values from a seed-like scalar and tile coordinates, and those values enter the first tensor-core computation. The best inference-level classification is therefore one-pass latent-conditioned generative rendering.

flowchart LR
  subgraph raw[Engine and temporal resources]
    color["Color / RGB"]
    mv["MVec + scale"]
    depth["Depth"]
    prev["dlssnr_prev_output"]
    control["ControlMask / UI / Backbuffer"]
  end

  submit["CG2R network-manager submission\nresource packing + CUBIN bindings"]
  pre["1H pre-block / input adapter\nRGB texture + feature packing"]
  rng["Internal Gaussian path\nseed-like field + tile coordinates\n3 generated FP16 lanes + 1.0"]
  e1["1H / 32\ntiled Swin"]
  e2["2H / 64\ntiled Swin"]
  e3["4H / 128\ntiled Swin"]
  e4["8H / 256\ntiled Swin"]
  core["16H split-Swin compound core\nFFN + fused QKVAttn + projections/pool\nfinal head: 512 -> 1024"]
  bridge["CCDecInputUpsample\n1024 -> 512\nmain + skip\noutput + reduce/overlap/conv-tap CBs"]
  d4["8H decoder\nupsample + skip"]
  d3["4H decoder\nupsample + skip"]
  d2["2H decoder\nupsample + skip"]
  d1["1H decoder\nupsample + skip"]
  post["1H post-block\n2 inputs -> CudaSurface\nmask / simple-blend variants"]
  output["Output\nplus history update"]

  subgraph islands["Registered ViT islands; active placement unresolved"]
    vit2d["ViT 2D compound block\nFFN + QKV + attention + projection"]
    vit1d["ViT 1D compound block\nFFN + QKV + attention + projection"]
    repack["2D <-> 1D repack"]
    vit2d -.-> vit1d -.-> repack
  end

  color --> submit
  mv --> submit
  depth --> submit
  prev --> submit
  submit --> pre
  rng -. generated inside pre-block .-> pre
  pre --> e1 --> e2 --> e3 --> e4 --> core --> bridge
  e4 -. skip .-> d4
  e3 -. skip .-> d3
  e2 -. skip .-> d2
  e1 -. skip .-> d1
  bridge --> d4 --> d3 --> d2 --> d1 --> post --> output
  control --> post
  output -. writes .-> prev

  classDef observed fill:#d9f2e6,stroke:#24734d,color:#123b29;
  classDef inferred fill:#fff1cc,stroke:#9a6b00,color:#4d3500;
  class submit,pre,rng,e1,e2,e3,e4,core,bridge,d4,d3,d2,d1,post observed;
  class raw,prev,output,vit2d,vit1d,repack inferred;

The solid path is a family-level skeleton, not a claim that each displayed level maps to exactly one serialized blockN

. The weight resource contains many repeated records and sub-block tensors. The dashed skip edges are structurally supported by the decoder kernels and wrapper contracts, but the exact merge operation is not yet separated from the fused code.

The ViT/1D blocks are shown as a separate registered island because their internal inventories are clear while the active descriptor sequence does not yet establish whether they sit before the 16H core, between the core and decoder, or in a parallel conditioning branch.

property recovered value confidence
target DLL NVIDIA NGX rel_310_8 , nvngx_dlssnr.dll
high
SHA-256 e16bcf15e16e13f527491cdf7845b2fe6521a738d8f7c9c721866a8496e1fc8e
high
learned-resource payload WEIGHTS_HT/1033 , 147,695,410 bytes
high
resource records exactly 153 parsed records across 71 logical block IDs (block0 –block70 )
high
embedded CUDA 15 fatbins and 231 public-inventory kernel names high
GPU target sm_120 / Blackwell only in this build
high
architecture hierarchical tiled Swin + split-Swin/ViT/1D compound runtime + decoder routing high for vocabulary; medium for exact topology
visible hierarchy 1H/32 → 2H/64 → 4H/128 → 8H/256, with a 16H/512 central core high for families; medium for placement/counts
central compound block six specialized layer types, including fused QKV-attention and a final width expansion high
decoder bridge CCDecInputUpsample , 1024 → 512, main + skip
high
parameter count approximately 148 million bytes of opaque learned-resource payload; exact independent scalar count remains unresolved medium
temporal path previous output bound into the network-manager submission and updated after evaluation high structurally; medium for exact slot
internal random path hash/coordinate-derived Gaussian-like FP16 lanes in the 1H pre-block high for mechanics; medium for semantics
inference style one forward submission; no visible timestep or denoising loop high

The public write-up's estimate of roughly 148 million FP8 parameters is credible as a rounded model-size figure, but the static resource parser now gives a more precise statement of what is actually measured. The WEIGHTS_HT/1033

resource is 147,695,410 bytes and contains exactly 153 named records. Their opaque payloads total 147,683,778 bytes; the remaining 11,632 bytes are the resource header, record metadata, padding, and footer. The record tags are raw storage markers (including several non-printable-looking values), not yet decoded E4M3/E5M3 labels. Some records also contain 2-byte or small auxiliary payloads. Consequently, the payload total is not yet an exact count of independent learned scalars.

What is now exact at the resource level is the named-record and block inventory: 71 logical IDs, with a near-symmetric 4/5/8-layer core pattern. That inventory materially constrains the topology even though the active descriptor order and storage decoding remain separate questions.

The resource parser follows the record framing and validates the resource-size field against the actual 147,695,410-byte blob. It recovers the public write-up's 153-record count, including the small block70.layer0.blend_scale

record that is easy to miss if only blockN.layerM.layer

names are counted. The resource contains 71 numeric block IDs (block0

through block70

). The most informative aggregate pattern is:

block IDs records per block payload bytes per block structural reading
block0 –block4
1 20–23 KiB edge/adapter-size records
block5 –block8
1 62–70 KiB small transition records
block9 –block14
1 197–230 KiB medium transition records
block15 –block22
1 689–820 KiB wide single-layer records
block23 –block29
4 1,968,192 repeated 4-layer compound group
block30
5 2,492,496 5-layer compound boundary group
block31 –block38
5 12,587,154 eight repeated central 5-layer groups
block39
1 525,312 central transition record
block40 –block47
4 1,968,192 repeated 4-layer compound group
block48 –block55
1 689–821 KiB wide single-layer records
block56 –block61
1 197–230 KiB medium transition records
block62 –block69
1 21–70 KiB small transition/adapter records
block70
2 21,810 2-byte blend_scale plus one layer record

The ranges above intentionally summarize similar payload sizes rather than assigning each ID to an encoder or decoder stage. The center's eight repeated five-record groups, flanked by repeated four-record groups, are strong structural corroboration for a large compound core surrounded by hierarchical transitions. They do not, by themselves, prove the runtime's active execution order or map one serialized block to one displayed Swin family.

The feature wrapper exposes DLSSNR.Color

, MVec

, and Depth

as core inputs, with optional ControlMask

, UI

, UIAlpha

, Backbuffer

, and BidirectionalDistortionField

resources. It also exposes controls including Intensity

, LocalToneStrength

, LocalStructureStrength

, SkinStructureStrength

, UseAutoMask

, UICorrection

, Enabled

, Reset

, and DepthInverted

.

The runtime maintains separate dlssnr_original_color

and dlssnr_network_output_scratch

surfaces in addition to dlssnr_prev_output

. The evaluator creates the scratch/original-color resources and invokes CG2RNetworkManager::Evaluate

(sub_180021bb0

). The manager reads the feature's previous-output field and sends it through the CUBIN resource-binding path. A later copy/update path writes the persistent history resource.

This is strong evidence for recurrent history feedback into the learned network submission rather than a history buffer used only by final blending. The exact history representation, reprojection, and insertion point remain open: the current evidence does not prove whether history is concatenated at the pre-block, injected deeper in the hierarchy, or transformed by a dedicated temporal kernel.

The CPU runtime registers 1H, 2H, 4H, and 8H fused Swin families. Representative CUDA modules use tensor-core MMA, explicit shared-memory staging, and variants for input/output views, chaining, waiting, tile synchronization, upsampling, and down/scale routing. The family names encode widths and head counts:

family visible width/head clue representative role
cc_tinlayout_fused_pre_block_swin_1h_32*
1H / 32-wide RGB-aware pre-block and first fused block
cc_tinlayout_fused_swin_1h_32*
1H / 32-wide regular fine-resolution local block
cc_tinlayout_fused_swin_2h_64_2*
2H / 64-wide hierarchical local block
cc_tinlayout_fused_swin_4h_128_4*
4H / 128-wide hierarchical local block
cc_tinlayout_fused_swin_8h_256_8*
8H / 256-wide hierarchical local block and decoder counterpart

The pre-block's template includes FusedSwin2d1HConfig<32,...>

and MpCubicSiluActivation

. The regular 1H body has 512 MMA operations and repeated rsqrt

operations; the representative 2H, 4H, and 8H bodies have approximately 512, 448, and 576 MMA operations respectively. Non-FP8 shared footprints grow from roughly 9 KiB to 17 KiB to 34 KiB across those wider families, consistent with larger local tiles and more parallel channels.

The decoder-specific variants expose _upsample

names and the host runtime has an explicit cc_tinlayout_upsample_skip_block

factory. This supports a U-Net-like hierarchy with encoder features retained for decoder-side routing, although the exact concat/add/gated merge is unresolved.

The host wrapper CCSplitSwin16HBlock

(sub_180040260

) recognizes six specialized layer types:

CCSplitSwin16HFfwd
CCSplitSwin16HFfwdProj
CCSplitSwin16HQKVAttn
CCSplitSwin16HProj
CCSplitSwin16HProjPool
CCSplitSwin16HFinalHead

The corresponding CUDA family contains separate FFN, FFN-projection, QKV, projection, projection+pool, and final-head entries. The host type is explicitly QKVAttn

, and there is no standalone split-Swin attention entry; the QKV family owns the fused middle operation. The base FFN entry contains 256 MMA operations, while the QKV entry contains 224 MMA operations plus warp-level reduction, normalization, and reciprocal work.

The final-head entry is a central feature transition, not the final image output head. Its template is:

Conv2d1x1Config<1024,512,...>

The first convolution template channel is the output width, as independently confirmed by the ViT FFN expand/contract pair. Therefore this layer expands 512 → 1024. The later decoder bridge uses Conv2d1x1Config<512,1024,...>

and contracts 1024 → 512. The host CCDecInputUpsample

constructor independently checks 0x400 → 0x200

and identifies its specialization as “1024->512 (dec5)”.

The host-side layer contracts make the split-Swin interface more specific than the kernel names alone. CCSplitSwin16HFfwd

is a one-input/one-output layer; CCSplitSwin16HFfwdProj

takes (skip, src)

and produces one output; CCSplitSwin16HQKVAttn

is one input/one output; CCSplitSwin16HProj

takes (src, skip)

and produces one output; CCSplitSwin16HProjPool

takes two inputs and produces two outputs; and CCSplitSwin16HFinalHead

is one input/one output. This is consistent with an internally routed compound block, rather than six independent top-level network stages.

The 2D ViT wrapper (sub_180041ca0

) recognizes five layer types:

CCVitFfnExpand
CCVitFfnContract
CCVitQKV
CCVitAttention
CCVitProjection

The 1D wrapper (sub_180043260

) recognizes corresponding 1D variants and explicitly checks for five layer descriptors. The CUDA family adds cc_vit_1d_repack_2d_to_1d

and cc_vit_1d_repack_1d_to_2d

entries.

Representative ViT QKV entries contain 192 MMA operations, consistent with three projections. Separate attention entries contain 128 MMA operations for QK-like work and weighted V accumulation. FFN expand/contract entries are separate; the 2D expand/contract templates are 1024 → 4096 and 4096 → 1024.

These blocks are high-confidence runtime capabilities and compound structures, but their active placement remains unresolved because the generic factory and constructor paths expose type dispatch rather than the serialized active network sequence.

The following contracts come from the host wrappers' validation strings and are more reliable than inferring interfaces from register allocation alone:

layer family inputs outputs / side buffers confidence
1H pre-block RGB texture input one output; _ds exposes (pool, swin)
high
1H/2H/4H/8H regular Swin one tensor; upsample variant additionally takes skip
one output; _ds exposes (pool, swin)
high
split-Swin FFN / FFN-projection one tensor / (skip, src)
one output high
split-Swin QKV-attention one tensor one output high
split-Swin projection (src, skip)
one output high
split-Swin projection+pool two tensors two outputs, likely pool plus transformed projection high for arity; medium for identity
split-Swin final head one tensor one tensor, 512 → 1024 high
2D ViT FFN expand one tensor one output high
2D ViT FFN contract (skip, src)
dst plus reduce CB
high
2D ViT QKV one tensor qry , key , val plus reduce CB
high
2D ViT attention qry , key , val
one output high
2D ViT projection (src, skip)
dst plus reduce CB
high
1D ViT FFN/QKV/attention/projection explicit token tensors and input/output CBs explicit Q/K/V and destination/reduction CBs high
decoder input upsample (main, skip)
{tensor, reduce_cb, overlap_cb, conv_tap}
high
post-block (main, enc0 skip)
one CudaSurface output
high

The ViT QKV contract is especially informative: the runtime really does materialize three logical projections (qry

, key

, val

) plus a reduction control buffer at the wrapper boundary, even though the GPU implementation fuses parts of their processing. The 1D wrapper likewise names explicit input/output CBs for the token path and exposes 2D↔1D repack kernels. This narrows the possible model interfaces substantially without claiming that these ViT islands are active in the shipping forward path.

The specialized CCDecInputUpsample

wrapper checks the channel transition 1024 → 512

and requires two inputs named (main, skip)

. It also validates a four-part output bundle:

{ tensor, reduce_cb, overlap_cb, conv_tap }

The matching normal decoder kernel is a no-activation Conv2d1x1Config<512,1024,...>

implementation with 64 MMA operations, async shared staging, and repeated vectorized output stores. The FP8 variant changes the storage tile configuration but preserves the same learned 1x1 bridge role. The PTX shows global reduction updates and a final integer release/store path indexed by tile/block coordinates, which explains why the wrapper exposes reduction and overlap control buffers in addition to the tensor result. The exact parameter-slot identity of reduce_cb

, overlap_cb

, and conv_tap

is not fully proven because the synchronization variants add pointer arguments, but the four-output contract itself is explicit.

The post kernel family is likewise more than a generic image store. The base CCTinlayoutFusedPostBlockSwin1HLayer

wrapper requires (main, enc0 skip)

and one CudaSurface

output. Its kernel variants include simple_blend

, control_mask

, and full_rect

forms, with corresponding FP8 variants. The PTX has a late texture-sampling/control path and a four-component surface write; the most defensible interpretation is three transformed output lanes plus a fourth alpha/control-like lane, followed by mode-dependent blending. The exact color-space transform and history-update equation remain open.

The CUDA templates name the activation MpCubicSiluActivation

. The emitted code does not call a conventional transcendental sigmoid. Instead, the representative low-precision path clamps the input and evaluates a cubic-like piecewise approximation:

t = clamp(x, -4, +4)
p = (-0.0559082) * abs(t) + 0.447266
a = t * p + 0.894531
y = x * a

The constants are rounded to half precision in the recovered path. This is a hardware-friendly SiLU-like nonlinearity that keeps the fused blocks tensor- core oriented.

The representative hierarchical Swin and split-QKV kernels contain FP32 reductions followed by rsqrt.approx.ftz.f32

, gain/bias-like multiplies, and reciprocal operations. This establishes an explicit variance- or energy-style normalization stage. The binary evidence available so far does not prove whether every instance subtracts a mean like LayerNorm or instead uses an RMS/L2-style normalization, so the exact normalization equation remains open.

The 2D and 1D ViT attention families are not simple unfused matrix-multiply wrappers. A representative cc_vit_1d_attention

path performs QK-like MMA, applies a half2 affine/clamp transform, constructs a fast positive exponential-like surrogate using bit manipulation, clamps the denominator to approximately 6.2e-5

, and accumulates V with reciprocal lanes. It is best described as softmax-like attention; the recovered code does not expose a canonical library softmax or ex2.approx

sequence.

The central split-Swin QKV family similarly fuses projection, reduction, normalization, reciprocal, and middle-operation work. Because the host type is CCSplitSwin16HQKVAttn

and there is no separate split-attention registration, the exact split attention/mixer equation cannot yet be isolated from its QKV kernel.

The embedded code targets Blackwell sm_120

and includes tensor-core FP8 conversion instructions such as UF2FP.SATFINITE.E4M3.F16

and F2FP.SATFINITE.E4M3.F16

; both E4M3 and E5M3 strings are present. The _fp8

variants are the explicit low-precision weight path. Other suffixes encode execution variants rather than new neural layers:

suffix/family likely role
_fp8
FP8 weight/dequantization path
_chained , _wait , _tilesync
dependency and tile-scheduling variants
_inpview , _outview
input/output view-layout variants
_upsample , *_ds
hierarchy transition or decoder variants
repack_2d_to_1d , repack_1d_to_2d
ViT 2D/1D layout conversion
control_mask , simple_blend , full_rect
post-block control/output modes

This naming matrix is one reason the public kernel count is much larger than the number of conceptual neural layers.

The strongest generative clue is inside the 1H pre-block rather than in a caller-visible input surface. The raw PTX for the representative pre entry contains a 264-byte parameter block, with a scalar at offset +200

and dimension-like values at +208

. The scalar is multiplied by -1640531527

, then combined with tile coordinates in a hash-like sequence.

The resulting code consumes four uniform-like values and executes the following recognizable pattern:

  • transform uniform values with lg2

,ln2

,-2

, andsqrt

; - use sin

andcos

with a2π

-like constant, as in a Box–Muller-style Gaussian construction; - convert generated values to FP16;

  • combine three generated lanes, a constant 1.0

, and four sampled texture lanes; - stage the resulting vector through shared memory before the first MMA.

The normal, FP8, and downsampled pre variants share this signature. No caller-provided noise surface, timestep, noise schedule, or loop over denoising steps was recovered. The safest interpretation is therefore an internal deterministic coordinate/seed-derived Gaussian-like conditioning path. Its exact semantic role—latent injection, stochastic feature augmentation, or a specialized preconditioning operation—remains unresolved.

The inference binary strongly supports a single-pass latent-conditioned generative renderer, but it does not contain the training metadata needed to name the objective. The current evidence ranks the interpretations as follows:

interpretation assessment from the binary
conventional multi-step diffusion sampler unlikely: no timestep/schedule state or denoising loop is visible in the analyzed path
one-step distilled diffusion, consistency, flow/meanflow, or DMD-like model strongly compatible with the one-pass latent-conditioned body and internal Gaussian-like path
GAN-style single-pass generator also compatible with the observed inference graph
deterministic neural renderer with a random-looking internal conditioner cannot be ruled out from implementation alone

DLSS5 appears to combine several performance choices that are mutually reinforcing:

  • Blackwell-only sm_120

kernels use tensor-core MMA paths and explicit FP8 conversion variants. - Windowed/tiled Swin-style operators keep attention local and fuse nearby projection, normalization, activation, and layout work.

  • The 16H compound core specializes its FFN, QKV-attention, projection, pool, and feature-transition stages instead of dispatching generic primitives.
  • Async dependency variants ( _chained

,_wait

,_tilesync

) let the runtime match kernel scheduling to the tile graph. - A single analyzed forward path avoids the cost of a conventional denoising loop.

  • Recurrent history feedback supplies temporal context without requiring a large sliding window of prior frames.

The current architecture is strong enough for a report, but several details would require live execution traces, more host-side state recovery, or weight interpretation:

  • the serialized active layer sequence, repetition counts, and exact placement of the 2D/1D ViT islands;
  • the meaning and lifetime of the +200

pre-block scalar and the precise mapping from generated lanes to network channels; - the pre-block's complete RGB/motion/depth packing and the exact shape of its internal Gaussian-like injection;

  • the exact mean-subtracting-vs-RMS normalization equation;
  • the full split-Swin QKV-attention/mixer equation;
  • whether decoder skip routing is add, concatenate, gated blend, or a fused equivalent, and the exact identity of the decoder reduce_cb

,overlap_cb

, andconv_tap

pointer slots; - the precise post-block color transform and output/history update semantics for each mask and blend mode;

  • the active descriptor's exact block sequence and the placement/repetition of the ViT/1D islands. The builder selects a runtime descriptor and dispatches through factories, but the serialized descriptor contents are not exposed as a simple string list in this build;
  • the training provenance: one-step diffusion distillation, consistency/flow matching, DMD-like regression, GAN, or another objective.

The following host-side procedures were the most useful anchors in Hopper:

location evidence recovered
sub_18003c280
family registration for pre/post 1H, hierarchical Swin, split-Swin, ViT, decoder bridge, and clear callback
sub_1800326c0
weight binding names: input_adapter_weight , weight1 , weight2 , ffn_cos_skip , qkv_weight , attn_scale , attn_bias , projection_weight , attn_cos_skip
sub_18003cc80
pre-block dispatch and the explicit “pre-block requires rgb input” guard
sub_18003fdc0
single-layer dispatcher for 1H/2H/4H/8H Swin, fused pre/post blocks, and decoder input upsample
sub_180040260
six-layer CCSplitSwin16HBlock compound wrapper
sub_180041ca0
five-layer 2D CCVitBlock wrapper
sub_180043260
five-layer 1D CCVit1DBlock wrapper and repack expectations
sub_180073bf0
decoder bridge dimension checks and “1024->512 (dec5)” specialization
sub_18001f570 range
active-network builder: descriptor selection, runtime dimension logging, consolidated aligned weight-heap sizing, and factory-driven network construction
sub_180021bb0
CG2RNetworkManager::Evaluate , resource setup, previous-output handling, and CUBIN binding path
tools/parse_dlss5_weight_map.py
validated 153-record WEIGHTS_HT map, 71 logical block IDs, payload totals, and raw storage tags

The CUDA-side evidence comes from the 15 embedded sm_120

fatbins, the representative PTX/SASS reductions, and the kernel-family naming matrix. Static analysis establishes capabilities and likely dataflow; it does not by itself prove the runtime's active descriptor sequence.

aspect DLSS 4.5 report DLSS5 finding
core organization compact local-transformer/U-Net-like body heterogeneous hierarchical Swin + 16H split-Swin core + registered ViT islands
attention/mixing local softmax attention followed by a softmax-free local mixer fused local Swin families, split QKV-attention, and ViT softmax-like attention
activation cubic SiLU approximation same named MpCubicSiluActivation family is present
normalization more fully characterized from the smaller graph rsqrt/gain-style normalization recovered; exact mean-vs-RMS form remains open
temporal path previous output/history participates in the graph recurrent previous-output/history path is again visible, with no evidence of a large sliding window
output path dynamic anisotropic Gaussian reconstruction filter was recovered post-block mask/simple-blend/full-rect families are present, but a comparable output filter has not been recovered
latent/random path no analogous internal Gaussian signature reported pre-block contains an internal Gaussian-like, coordinate/seed-derived path
inference interpretation deterministic neural upscaler with specialized local mixing single-pass latent-conditioned generative neural renderer
training claim architecture analysis did not require a training-objective claim one-step distilled diffusion is plausible, but not distinguishable from GAN/flow/consistency alternatives from the DLL alone

The comparison is intentionally asymmetric: DLSS4.5's smaller model was amenable to a more complete graph reconstruction, while the DLSS5 artifact exposes a larger, more heterogeneous runtime with several compound islands whose active ordering is not serialized in the public evidence we have.

DLSS5 is best described as a recurrent, single-pass, latent-conditioned generative neural renderer. Its body combines a 1H input adapter, hierarchical 1H/2H/4H/8H tiled Swin stages, a 16H split-Swin compound core with a 512 → 1024 feature expansion, a 1024 → 512 decoder bridge, decoder-side upsample/ skip blocks, and a 1H mask/blend post path. ViT and 1D transformer blocks are also registered as specialized compound islands, although their active placement is not yet proven.

The internal coordinate/seed-derived Gaussian-like path and the absence of a runtime timestep or denoising loop make one-step generative inference a strong architectural interpretation. They do not, however, identify whether NVIDIA trained it with meanflow, consistency distillation, DMD, GAN losses, or some other teacher/student objective. That distinction belongs in the “likely training neighborhood,” not in the list of facts directly recovered from the binary.

── more in #machine-learning 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-is-the-dlss5-mo…] indexed:0 read:20min 2026-08-30 · —