What is the DLSS5 model architecture? A developer reverse-engineered NVIDIA's DLSS5 binary to infer its model architecture, revealing a heterogeneous, hierarchical, recurrent, single-pass neural renderer with tiled Swin transformer families and a compound 16H split-Swin core. The analysis, published in a gist, indicates the runtime exposes no timestep or denoising loop, classifying the model as one-pass latent-conditioned generative rendering. Let's attempt to figure out the model architecture from DLSS5's binary See https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee https://gist.github.com/madebyollin/f87b506b2779c2aeae5e44c81ef0fdee for context. Warning The remainder of this gist was authored by Codex 5.6 The strongest current model is a heterogeneous, hierarchical, recurrent, single-pass neural renderer . Its visible body combines tiled 1H/2H/4H/8H Swin families, a compound 16H split-Swin core, 2D and 1D ViT compound blocks, specialized 1x1 width transitions, and a decoder with upsample/skip routing. The most important difference from a conventional diffusion implementation is that the runtime exposes no timestep, noise schedule, or loop over denoising steps. However, the first 1H pre-block internally synthesizes Gaussian-like values from a seed-like scalar and tile coordinates, and those values enter the first tensor-core computation. The best inference-level classification is therefore one-pass latent-conditioned generative rendering. flowchart LR subgraph raw Engine and temporal resources color "Color / RGB" mv "MVec + scale" depth "Depth" prev "dlssnr prev output" control "ControlMask / UI / Backbuffer" end submit "CG2R network-manager submission\nresource packing + CUBIN bindings" pre "1H pre-block / input adapter\nRGB texture + feature packing" rng "Internal Gaussian path\nseed-like field + tile coordinates\n3 generated FP16 lanes + 1.0" e1 "1H / 32\ntiled Swin" e2 "2H / 64\ntiled Swin" e3 "4H / 128\ntiled Swin" e4 "8H / 256\ntiled Swin" core "16H split-Swin compound core\nFFN + fused QKVAttn + projections/pool\nfinal head: 512 - 1024" bridge "CCDecInputUpsample\n1024 - 512\nmain + skip\noutput + reduce/overlap/conv-tap CBs" d4 "8H decoder\nupsample + skip" d3 "4H decoder\nupsample + skip" d2 "2H decoder\nupsample + skip" d1 "1H decoder\nupsample + skip" post "1H post-block\n2 inputs - CudaSurface\nmask / simple-blend variants" output "Output\nplus history update" subgraph islands "Registered ViT islands; active placement unresolved" vit2d "ViT 2D compound block\nFFN + QKV + attention + projection" vit1d "ViT 1D compound block\nFFN + QKV + attention + projection" repack "2D <- 1D repack" vit2d -.- vit1d -.- repack end color -- submit mv -- submit depth -- submit prev -- submit submit -- pre rng -. generated inside pre-block .- pre pre -- e1 -- e2 -- e3 -- e4 -- core -- bridge e4 -. skip .- d4 e3 -. skip .- d3 e2 -. skip .- d2 e1 -. skip .- d1 bridge -- d4 -- d3 -- d2 -- d1 -- post -- output control -- post output -. writes .- prev classDef observed fill: d9f2e6,stroke: 24734d,color: 123b29; classDef inferred fill: fff1cc,stroke: 9a6b00,color: 4d3500; class submit,pre,rng,e1,e2,e3,e4,core,bridge,d4,d3,d2,d1,post observed; class raw,prev,output,vit2d,vit1d,repack inferred; The solid path is a family-level skeleton, not a claim that each displayed level maps to exactly one serialized blockN . The weight resource contains many repeated records and sub-block tensors. The dashed skip edges are structurally supported by the decoder kernels and wrapper contracts, but the exact merge operation is not yet separated from the fused code. The ViT/1D blocks are shown as a separate registered island because their internal inventories are clear while the active descriptor sequence does not yet establish whether they sit before the 16H core, between the core and decoder, or in a parallel conditioning branch. | property | recovered value | confidence | |---|---|---| | target DLL | NVIDIA NGX rel 310 8 , nvngx dlssnr.dll | high | | SHA-256 | e16bcf15e16e13f527491cdf7845b2fe6521a738d8f7c9c721866a8496e1fc8e | high | | learned-resource payload | WEIGHTS HT/1033 , 147,695,410 bytes | high | | resource records | exactly 153 parsed records across 71 logical block IDs block0 – block70 | high | | embedded CUDA | 15 fatbins and 231 public-inventory kernel names | high | | GPU target | sm 120 / Blackwell only in this build | high | | architecture | hierarchical tiled Swin + split-Swin/ViT/1D compound runtime + decoder routing | high for vocabulary; medium for exact topology | | visible hierarchy | 1H/32 → 2H/64 → 4H/128 → 8H/256, with a 16H/512 central core | high for families; medium for placement/counts | | central compound block | six specialized layer types, including fused QKV-attention and a final width expansion | high | | decoder bridge | CCDecInputUpsample , 1024 → 512, main + skip | high | | parameter count | approximately 148 million bytes of opaque learned-resource payload; exact independent scalar count remains unresolved | medium | | temporal path | previous output bound into the network-manager submission and updated after evaluation | high structurally; medium for exact slot | | internal random path | hash/coordinate-derived Gaussian-like FP16 lanes in the 1H pre-block | high for mechanics; medium for semantics | | inference style | one forward submission; no visible timestep or denoising loop | high | The public write-up's estimate of roughly 148 million FP8 parameters is credible as a rounded model-size figure, but the static resource parser now gives a more precise statement of what is actually measured. The WEIGHTS HT/1033 resource is 147,695,410 bytes and contains exactly 153 named records. Their opaque payloads total 147,683,778 bytes; the remaining 11,632 bytes are the resource header, record metadata, padding, and footer. The record tags are raw storage markers including several non-printable-looking values , not yet decoded E4M3/E5M3 labels. Some records also contain 2-byte or small auxiliary payloads. Consequently, the payload total is not yet an exact count of independent learned scalars. What is now exact at the resource level is the named-record and block inventory: 71 logical IDs, with a near-symmetric 4/5/8-layer core pattern. That inventory materially constrains the topology even though the active descriptor order and storage decoding remain separate questions. The resource parser follows the record framing and validates the resource-size field against the actual 147,695,410-byte blob. It recovers the public write-up's 153-record count, including the small block70.layer0.blend scale record that is easy to miss if only blockN.layerM.layer names are counted. The resource contains 71 numeric block IDs block0 through block70 . The most informative aggregate pattern is: | block IDs | records per block | payload bytes per block | structural reading | |---|---|---|---| block0 – block4 | 1 | 20–23 KiB | edge/adapter-size records | block5 – block8 | 1 | 62–70 KiB | small transition records | block9 – block14 | 1 | 197–230 KiB | medium transition records | block15 – block22 | 1 | 689–820 KiB | wide single-layer records | block23 – block29 | 4 | 1,968,192 | repeated 4-layer compound group | block30 | 5 | 2,492,496 | 5-layer compound boundary group | block31 – block38 | 5 | 12,587,154 | eight repeated central 5-layer groups | block39 | 1 | 525,312 | central transition record | block40 – block47 | 4 | 1,968,192 | repeated 4-layer compound group | block48 – block55 | 1 | 689–821 KiB | wide single-layer records | block56 – block61 | 1 | 197–230 KiB | medium transition records | block62 – block69 | 1 | 21–70 KiB | small transition/adapter records | block70 | 2 | 21,810 | 2-byte blend scale plus one layer record | The ranges above intentionally summarize similar payload sizes rather than assigning each ID to an encoder or decoder stage. The center's eight repeated five-record groups, flanked by repeated four-record groups, are strong structural corroboration for a large compound core surrounded by hierarchical transitions. They do not, by themselves, prove the runtime's active execution order or map one serialized block to one displayed Swin family. The feature wrapper exposes DLSSNR.Color , MVec , and Depth as core inputs, with optional ControlMask , UI , UIAlpha , Backbuffer , and BidirectionalDistortionField resources. It also exposes controls including Intensity , LocalToneStrength , LocalStructureStrength , SkinStructureStrength , UseAutoMask , UICorrection , Enabled , Reset , and DepthInverted . The runtime maintains separate dlssnr original color and dlssnr network output scratch surfaces in addition to dlssnr prev output . The evaluator creates the scratch/original-color resources and invokes CG2RNetworkManager::Evaluate sub 180021bb0 . The manager reads the feature's previous-output field and sends it through the CUBIN resource-binding path. A later copy/update path writes the persistent history resource. This is strong evidence for recurrent history feedback into the learned network submission rather than a history buffer used only by final blending. The exact history representation, reprojection, and insertion point remain open: the current evidence does not prove whether history is concatenated at the pre-block, injected deeper in the hierarchy, or transformed by a dedicated temporal kernel. The CPU runtime registers 1H, 2H, 4H, and 8H fused Swin families. Representative CUDA modules use tensor-core MMA, explicit shared-memory staging, and variants for input/output views, chaining, waiting, tile synchronization, upsampling, and down/scale routing. The family names encode widths and head counts: | family | visible width/head clue | representative role | |---|---|---| cc tinlayout fused pre block swin 1h 32 | 1H / 32-wide | RGB-aware pre-block and first fused block | cc tinlayout fused swin 1h 32 | 1H / 32-wide | regular fine-resolution local block | cc tinlayout fused swin 2h 64 2 | 2H / 64-wide | hierarchical local block | cc tinlayout fused swin 4h 128 4 | 4H / 128-wide | hierarchical local block | cc tinlayout fused swin 8h 256 8 | 8H / 256-wide | hierarchical local block and decoder counterpart | The pre-block's template includes FusedSwin2d1HConfig<32,... and MpCubicSiluActivation . The regular 1H body has 512 MMA operations and repeated rsqrt operations; the representative 2H, 4H, and 8H bodies have approximately 512, 448, and 576 MMA operations respectively. Non-FP8 shared footprints grow from roughly 9 KiB to 17 KiB to 34 KiB across those wider families, consistent with larger local tiles and more parallel channels. The decoder-specific variants expose upsample names and the host runtime has an explicit cc tinlayout upsample skip block factory. This supports a U-Net-like hierarchy with encoder features retained for decoder-side routing, although the exact concat/add/gated merge is unresolved. The host wrapper CCSplitSwin16HBlock sub 180040260 recognizes six specialized layer types: CCSplitSwin16HFfwd CCSplitSwin16HFfwdProj CCSplitSwin16HQKVAttn CCSplitSwin16HProj CCSplitSwin16HProjPool CCSplitSwin16HFinalHead The corresponding CUDA family contains separate FFN, FFN-projection, QKV, projection, projection+pool, and final-head entries. The host type is explicitly QKVAttn , and there is no standalone split-Swin attention entry; the QKV family owns the fused middle operation. The base FFN entry contains 256 MMA operations, while the QKV entry contains 224 MMA operations plus warp-level reduction, normalization, and reciprocal work. The final-head entry is a central feature transition, not the final image output head. Its template is: Conv2d1x1Config<1024,512,... The first convolution template channel is the output width, as independently confirmed by the ViT FFN expand/contract pair. Therefore this layer expands 512 → 1024. The later decoder bridge uses Conv2d1x1Config<512,1024,... and contracts 1024 → 512. The host CCDecInputUpsample constructor independently checks 0x400 → 0x200 and identifies its specialization as “1024- 512 dec5 ”. The host-side layer contracts make the split-Swin interface more specific than the kernel names alone. CCSplitSwin16HFfwd is a one-input/one-output layer; CCSplitSwin16HFfwdProj takes skip, src and produces one output; CCSplitSwin16HQKVAttn is one input/one output; CCSplitSwin16HProj takes src, skip and produces one output; CCSplitSwin16HProjPool takes two inputs and produces two outputs; and CCSplitSwin16HFinalHead is one input/one output. This is consistent with an internally routed compound block, rather than six independent top-level network stages. The 2D ViT wrapper sub 180041ca0 recognizes five layer types: CCVitFfnExpand CCVitFfnContract CCVitQKV CCVitAttention CCVitProjection The 1D wrapper sub 180043260 recognizes corresponding 1D variants and explicitly checks for five layer descriptors. The CUDA family adds cc vit 1d repack 2d to 1d and cc vit 1d repack 1d to 2d entries. Representative ViT QKV entries contain 192 MMA operations, consistent with three projections. Separate attention entries contain 128 MMA operations for QK-like work and weighted V accumulation. FFN expand/contract entries are separate; the 2D expand/contract templates are 1024 → 4096 and 4096 → 1024. These blocks are high-confidence runtime capabilities and compound structures, but their active placement remains unresolved because the generic factory and constructor paths expose type dispatch rather than the serialized active network sequence. The following contracts come from the host wrappers' validation strings and are more reliable than inferring interfaces from register allocation alone: | layer family | inputs | outputs / side buffers | confidence | |---|---|---|---| | 1H pre-block | RGB texture input | one output; ds exposes pool, swin | high | | 1H/2H/4H/8H regular Swin | one tensor; upsample variant additionally takes skip | one output; ds exposes pool, swin | high | | split-Swin FFN / FFN-projection | one tensor / skip, src | one output | high | | split-Swin QKV-attention | one tensor | one output | high | | split-Swin projection | src, skip | one output | high | | split-Swin projection+pool | two tensors | two outputs, likely pool plus transformed projection | high for arity; medium for identity | | split-Swin final head | one tensor | one tensor, 512 → 1024 | high | | 2D ViT FFN expand | one tensor | one output | high | | 2D ViT FFN contract | skip, src | dst plus reduce CB | high | | 2D ViT QKV | one tensor | qry , key , val plus reduce CB | high | | 2D ViT attention | qry , key , val | one output | high | | 2D ViT projection | src, skip | dst plus reduce CB | high | | 1D ViT FFN/QKV/attention/projection | explicit token tensors and input/output CBs | explicit Q/K/V and destination/reduction CBs | high | | decoder input upsample | main, skip | {tensor, reduce cb, overlap cb, conv tap} | high | | post-block | main, enc0 skip | one CudaSurface output | high | The ViT QKV contract is especially informative: the runtime really does materialize three logical projections qry , key , val plus a reduction control buffer at the wrapper boundary, even though the GPU implementation fuses parts of their processing. The 1D wrapper likewise names explicit input/output CBs for the token path and exposes 2D↔1D repack kernels. This narrows the possible model interfaces substantially without claiming that these ViT islands are active in the shipping forward path. The specialized CCDecInputUpsample wrapper checks the channel transition 1024 → 512 and requires two inputs named main, skip . It also validates a four-part output bundle: { tensor, reduce cb, overlap cb, conv tap } The matching normal decoder kernel is a no-activation Conv2d1x1Config<512,1024,... implementation with 64 MMA operations, async shared staging, and repeated vectorized output stores. The FP8 variant changes the storage tile configuration but preserves the same learned 1x1 bridge role. The PTX shows global reduction updates and a final integer release/store path indexed by tile/block coordinates, which explains why the wrapper exposes reduction and overlap control buffers in addition to the tensor result. The exact parameter-slot identity of reduce cb , overlap cb , and conv tap is not fully proven because the synchronization variants add pointer arguments, but the four-output contract itself is explicit. The post kernel family is likewise more than a generic image store. The base CCTinlayoutFusedPostBlockSwin1HLayer wrapper requires main, enc0 skip and one CudaSurface output. Its kernel variants include simple blend , control mask , and full rect forms, with corresponding FP8 variants. The PTX has a late texture-sampling/control path and a four-component surface write; the most defensible interpretation is three transformed output lanes plus a fourth alpha/control-like lane, followed by mode-dependent blending. The exact color-space transform and history-update equation remain open. The CUDA templates name the activation MpCubicSiluActivation . The emitted code does not call a conventional transcendental sigmoid. Instead, the representative low-precision path clamps the input and evaluates a cubic-like piecewise approximation: t = clamp x, -4, +4 p = -0.0559082 abs t + 0.447266 a = t p + 0.894531 y = x a The constants are rounded to half precision in the recovered path. This is a hardware-friendly SiLU-like nonlinearity that keeps the fused blocks tensor- core oriented. The representative hierarchical Swin and split-QKV kernels contain FP32 reductions followed by rsqrt.approx.ftz.f32 , gain/bias-like multiplies, and reciprocal operations. This establishes an explicit variance- or energy-style normalization stage. The binary evidence available so far does not prove whether every instance subtracts a mean like LayerNorm or instead uses an RMS/L2-style normalization, so the exact normalization equation remains open. The 2D and 1D ViT attention families are not simple unfused matrix-multiply wrappers. A representative cc vit 1d attention path performs QK-like MMA, applies a half2 affine/clamp transform, constructs a fast positive exponential-like surrogate using bit manipulation, clamps the denominator to approximately 6.2e-5 , and accumulates V with reciprocal lanes. It is best described as softmax-like attention; the recovered code does not expose a canonical library softmax or ex2.approx sequence. The central split-Swin QKV family similarly fuses projection, reduction, normalization, reciprocal, and middle-operation work. Because the host type is CCSplitSwin16HQKVAttn and there is no separate split-attention registration, the exact split attention/mixer equation cannot yet be isolated from its QKV kernel. The embedded code targets Blackwell sm 120 and includes tensor-core FP8 conversion instructions such as UF2FP.SATFINITE.E4M3.F16 and F2FP.SATFINITE.E4M3.F16 ; both E4M3 and E5M3 strings are present. The fp8 variants are the explicit low-precision weight path. Other suffixes encode execution variants rather than new neural layers: | suffix/family | likely role | |---|---| fp8 | FP8 weight/dequantization path | chained , wait , tilesync | dependency and tile-scheduling variants | inpview , outview | input/output view-layout variants | upsample , ds | hierarchy transition or decoder variants | repack 2d to 1d , repack 1d to 2d | ViT 2D/1D layout conversion | control mask , simple blend , full rect | post-block control/output modes | This naming matrix is one reason the public kernel count is much larger than the number of conceptual neural layers. The strongest generative clue is inside the 1H pre-block rather than in a caller-visible input surface. The raw PTX for the representative pre entry contains a 264-byte parameter block, with a scalar at offset +200 and dimension-like values at +208 . The scalar is multiplied by -1640531527 , then combined with tile coordinates in a hash-like sequence. The resulting code consumes four uniform-like values and executes the following recognizable pattern: - transform uniform values with lg2 , ln2 , -2 , and sqrt ; - use sin and cos with a 2π -like constant, as in a Box–Muller-style Gaussian construction; - convert generated values to FP16; - combine three generated lanes, a constant 1.0 , and four sampled texture lanes; - stage the resulting vector through shared memory before the first MMA. The normal, FP8, and downsampled pre variants share this signature. No caller-provided noise surface, timestep, noise schedule, or loop over denoising steps was recovered. The safest interpretation is therefore an internal deterministic coordinate/seed-derived Gaussian-like conditioning path. Its exact semantic role—latent injection, stochastic feature augmentation, or a specialized preconditioning operation—remains unresolved. The inference binary strongly supports a single-pass latent-conditioned generative renderer, but it does not contain the training metadata needed to name the objective. The current evidence ranks the interpretations as follows: | interpretation | assessment from the binary | |---|---| | conventional multi-step diffusion sampler | unlikely: no timestep/schedule state or denoising loop is visible in the analyzed path | | one-step distilled diffusion, consistency, flow/meanflow, or DMD-like model | strongly compatible with the one-pass latent-conditioned body and internal Gaussian-like path | | GAN-style single-pass generator | also compatible with the observed inference graph | | deterministic neural renderer with a random-looking internal conditioner | cannot be ruled out from implementation alone | DLSS5 appears to combine several performance choices that are mutually reinforcing: - Blackwell-only sm 120 kernels use tensor-core MMA paths and explicit FP8 conversion variants. - Windowed/tiled Swin-style operators keep attention local and fuse nearby projection, normalization, activation, and layout work. - The 16H compound core specializes its FFN, QKV-attention, projection, pool, and feature-transition stages instead of dispatching generic primitives. - Async dependency variants chained , wait , tilesync let the runtime match kernel scheduling to the tile graph. - A single analyzed forward path avoids the cost of a conventional denoising loop. - Recurrent history feedback supplies temporal context without requiring a large sliding window of prior frames. The current architecture is strong enough for a report, but several details would require live execution traces, more host-side state recovery, or weight interpretation: - the serialized active layer sequence, repetition counts, and exact placement of the 2D/1D ViT islands; - the meaning and lifetime of the +200 pre-block scalar and the precise mapping from generated lanes to network channels; - the pre-block's complete RGB/motion/depth packing and the exact shape of its internal Gaussian-like injection; - the exact mean-subtracting-vs-RMS normalization equation; - the full split-Swin QKV-attention/mixer equation; - whether decoder skip routing is add, concatenate, gated blend, or a fused equivalent, and the exact identity of the decoder reduce cb , overlap cb , and conv tap pointer slots; - the precise post-block color transform and output/history update semantics for each mask and blend mode; - the active descriptor's exact block sequence and the placement/repetition of the ViT/1D islands. The builder selects a runtime descriptor and dispatches through factories, but the serialized descriptor contents are not exposed as a simple string list in this build; - the training provenance: one-step diffusion distillation, consistency/flow matching, DMD-like regression, GAN, or another objective. The following host-side procedures were the most useful anchors in Hopper: | location | evidence recovered | |---|---| sub 18003c280 | family registration for pre/post 1H, hierarchical Swin, split-Swin, ViT, decoder bridge, and clear callback | sub 1800326c0 | weight binding names: input adapter weight , weight1 , weight2 , ffn cos skip , qkv weight , attn scale , attn bias , projection weight , attn cos skip | sub 18003cc80 | pre-block dispatch and the explicit “pre-block requires rgb input” guard | sub 18003fdc0 | single-layer dispatcher for 1H/2H/4H/8H Swin, fused pre/post blocks, and decoder input upsample | sub 180040260 | six-layer CCSplitSwin16HBlock compound wrapper | sub 180041ca0 | five-layer 2D CCVitBlock wrapper | sub 180043260 | five-layer 1D CCVit1DBlock wrapper and repack expectations | sub 180073bf0 | decoder bridge dimension checks and “1024- 512 dec5 ” specialization | sub 18001f570 range | active-network builder: descriptor selection, runtime dimension logging, consolidated aligned weight-heap sizing, and factory-driven network construction | sub 180021bb0 | CG2RNetworkManager::Evaluate , resource setup, previous-output handling, and CUBIN binding path | tools/parse dlss5 weight map.py | validated 153-record WEIGHTS HT map, 71 logical block IDs, payload totals, and raw storage tags | The CUDA-side evidence comes from the 15 embedded sm 120 fatbins, the representative PTX/SASS reductions, and the kernel-family naming matrix. Static analysis establishes capabilities and likely dataflow; it does not by itself prove the runtime's active descriptor sequence. | aspect | DLSS 4.5 report | DLSS5 finding | |---|---|---| | core organization | compact local-transformer/U-Net-like body | heterogeneous hierarchical Swin + 16H split-Swin core + registered ViT islands | | attention/mixing | local softmax attention followed by a softmax-free local mixer | fused local Swin families, split QKV-attention, and ViT softmax-like attention | | activation | cubic SiLU approximation | same named MpCubicSiluActivation family is present | | normalization | more fully characterized from the smaller graph | rsqrt/gain-style normalization recovered; exact mean-vs-RMS form remains open | | temporal path | previous output/history participates in the graph | recurrent previous-output/history path is again visible, with no evidence of a large sliding window | | output path | dynamic anisotropic Gaussian reconstruction filter was recovered | post-block mask/simple-blend/full-rect families are present, but a comparable output filter has not been recovered | | latent/random path | no analogous internal Gaussian signature reported | pre-block contains an internal Gaussian-like, coordinate/seed-derived path | | inference interpretation | deterministic neural upscaler with specialized local mixing | single-pass latent-conditioned generative neural renderer | | training claim | architecture analysis did not require a training-objective claim | one-step distilled diffusion is plausible, but not distinguishable from GAN/flow/consistency alternatives from the DLL alone | The comparison is intentionally asymmetric: DLSS4.5's smaller model was amenable to a more complete graph reconstruction, while the DLSS5 artifact exposes a larger, more heterogeneous runtime with several compound islands whose active ordering is not serialized in the public evidence we have. DLSS5 is best described as a recurrent, single-pass, latent-conditioned generative neural renderer. Its body combines a 1H input adapter, hierarchical 1H/2H/4H/8H tiled Swin stages, a 16H split-Swin compound core with a 512 → 1024 feature expansion, a 1024 → 512 decoder bridge, decoder-side upsample/ skip blocks, and a 1H mask/blend post path. ViT and 1D transformer blocks are also registered as specialized compound islands, although their active placement is not yet proven. The internal coordinate/seed-derived Gaussian-like path and the absence of a runtime timestep or denoising loop make one-step generative inference a strong architectural interpretation. They do not, however, identify whether NVIDIA trained it with meanflow, consistency distillation, DMD, GAN losses, or some other teacher/student objective. That distinction belongs in the “likely training neighborhood,” not in the list of facts directly recovered from the binary.