{"slug": "threejs-port-of-opendlss-nr", "title": "ThreeJS Port of OpenDLSS-NR", "summary": "Developer Ben Houston released three-dlss-nr, a Three.js port of maanHimself's OpenDLSS-NR that runs the 71-block DLSS 5 neural rendering network as TSL compute kernels inside a WebGPURenderer with no CPU readback. The port tracks upstream commit 9d08f41 and is byte-identical to the reference WebGPU implementation across all 83 tensors and 451 compute dispatches at 64x64 and 512x512, verified in Chrome on an RTX 3060 Ti. It ships with synthetic weights only, so the demo output is not meaningful, and it is not affiliated with NVIDIA and includes no NVIDIA weights.", "body_md": "**A port to Three.js (TSL / WebGPU) of [OpenDLSS-NR](https://github.com/maanHimself/OpenDLSS-NR) by\n[maan](https://github.com/maanHimself).**\n\nOpenDLSS-NR is an open-source reimplementation of the network behind NVIDIA's DLSS 5 Neural Rendering. All credit\nfor that reimplementation, its numerics and its documentation belongs to OpenDLSS-NR. This repository ports it to\nthree.js and tracks upstream commit [`9d08f41`](https://github.com/maanHimself/OpenDLSS-NR/tree/9d08f41), included\nas the [`reference/OpenDLSS-NR`](https://github.com/bhouston/three-dlss-nr/blob/main/reference/OpenDLSS-NR) submodule. Not affiliated with NVIDIA, and no NVIDIA weights\nare included ([details](#license-and-notices)).\n\n**The DLSS 5 neural rendering network, as three.js compute nodes, byte for byte.** three-dlss-nr runs the 71-block\nOpenDLSS-NR network as [TSL](https://threejs.org/docs/#api/en/nodes/TSL) compute kernels inside a `WebGPURenderer`,\nfed straight from your three.js scene with no CPU readback. Every intermediate tensor is byte-identical to the\nreference WebGPU port.\n\n**Synthetic weights: output is not meaningful.** The demo in split view on the TSL backend, NR off on the left and\nNR on on the right. Synthetic weights have the right layout and run the network's real arithmetic, but the image they\nproduce means nothing, hence the red field ([why](#why-there-are-no-real-weights-and-what-the-synthetic-ones-are-for)).\nWith a real model directory the right half is the re-rendered head. Head: Lee Perry-Smith, CC BY 3.0\n([credits](#credits)).\n\nTry it live at **[three-dlss-nr.ben3d.ca](https://three-dlss-nr.ben3d.ca)**.\n\nFrom [upstream's README](https://github.com/maanHimself/OpenDLSS-NR/tree/9d08f41#the-network): it is a generative\nneural rendering network (NVIDIA's term). It **re-renders the frame the engine already drew**, generating detail from\ninjected noise and adjusting tone, structure and skin under a style setting. Input and output are the same\nresolution; **it is not an upscaler**.\n\n- A U-net of shifted-window (Swin) transformer blocks with a global ViT at the bottom: 71 blocks over six pooling levels, FP8 (E4M3) activations with FP16 accumulation, 141 MiB of weights.\n- In: one rendered frame (a low dynamic range proxy of it, three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars). Out: four f32 channels per pixel, an RGB residual and one temporal-blend logit.\n- 451 compute dispatches per frame in the WebGPU formulation.\n\nNVIDIA describes the model in its report,\n[DLSS 5: Generative Neural Rendering](https://research.nvidia.com/labs/adlr/DLSS5/files/DLSS5_Report.pdf)\n([project page](https://research.nvidia.com/labs/adlr/DLSS5/)). Upstream's\n[`docs/network.md`](https://github.com/maanHimself/OpenDLSS-NR/blob/9d08f41/docs/network.md) is the graph in full.\n\nThe spec for every kernel is upstream's WebGPU port (`ports/browser-webgpu`), and the port is checked against it on\nthe same device, the same inputs and the same weights:\n\n- **Every kernel family** (FP8 and f16 GEMMs, window attention, the global ViT, the elementwise ops, preprocess, and\nthe frame kernels that build the input features and compose the output) is byte-identical to the reference.\n- **The whole network** (451 dispatches) is byte-identical at 64x64 and at 512x512 on synthetic weights: all 79 block\nboundaries, the three post tensors and the head, 83 of 83 tensors, with repeat runs identical. Checked in real\nChrome on an RTX 3060 Ti (D3D12 + DXC), with both networks in one page on one device.\n- **The numerics** (E4M3 encode and decode, f16 rounding, the SiLU and exponent tables, FP8 and f16 dot products) are\nchecked on the GPU against all seven sections of the reference's numerics fixture, which upstream's Vulkan\nimplementation produced: exhaustively over every half and every byte where the domain allows it.\n- **CI** (Linux, Mesa lavapipe) runs the whole network in Node and gates it against golden digests recorded from the\nChrome run. It cannot compare against the reference directly: llvmpipe folds the`f32(f16(x))` round trip into`x` ,\nso the reference's f16 roundings silently do not happen there. The TSL port never relies on f16 hardware. It rounds\non f32 bit patterns, so it produces the same bytes on lavapipe as on a real GPU.\n\n**See it side by side:** [three-dlss-nr.ben3d.ca/parity/](https://three-dlss-nr.ben3d.ca/parity/) shows the\nreference running standalone, the TSL port and the reference shim on the same frames (the head scan and two\nprocedural scenes, at 256x256 and 512x512): the composed output, the head and four block boundaries, all byte-identical,\nwith every raw tensor compared and per-scene GPU times. It is built with [fidelity-kit](https://fidelity-kit.ben3d.ca)\nfrom [`packages/fidelity-suite`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/fidelity-suite), and CI checks its committed verdicts.\n\nA consequence: **the TSL port does not need `shader-f16`**, while the reference does. How the kernels are written and\ntested is in [`src/README-internals.md`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/three-dlss-nr/src/README-internals.md).\n\n```\nnpm install three-dlss-nr three\n```\n\n`DlssNrPass` puts the network behind a three.js scene. It renders the scene into an RGBA16F target (linear colour\nplus three's `velocity`), builds the network's input features on the GPU, runs the network, composes the result with\nthe temporal history, and draws NR off, NR on or a split view onto the canvas.\n\n``` js\nimport { backendBuilder, createNRRenderer, DlssNrPass, NRModel, tslBackend } from 'three-dlss-nr';\n\nconst { renderer } = await createNRRenderer(); // a WebGPURenderer on a device with the limits the network needs\ndocument.body.appendChild(renderer.domElement);\n\n// Your model directory: manifest.json + model/stages/* (see the weights section below).\nconst model = await NRModel.load('https://example.com/my-model');\n\nconst pass = new DlssNrPass({ renderer, scene, camera, width: 960, height: 540, view: 'split' });\nawait pass.setNetwork(backendBuilder(tslBackend, model)); // compiles the kernels; setSize() rebuilds\npass.setSettings({ intensity: 1, style: 1 }); // style: 0 none, 1 cinematic, 2 natural\n\nrenderer.setAnimationLoop(async () => {\n  const stats = await pass.render(); // null while the previous frame is still on the GPU\n  // stats?.network: { method, gpuMilliseconds, wallMilliseconds }\n});\n```\n\nLoad the model once and share it: resizes and network rebuilds reuse it instead of re-reading 141 MiB. Call\n`pass.resetHistory()` on a camera cut. The internal size is capped at 1280x720.\n\nTo run the network on your own inputs instead of a scene, use a backend directly:\n\n``` js\nconst network = await tslBackend.create({ renderer, model, width: 512, height: 512, onProgress: console.log });\nnetwork.writeFeatures(features); // Float32Array, fullRows x 16 lanes (or write network.features from a compute pass)\nconst timing = await network.run({ timing: true });\nconst head = await network.readHead(); // Float32Array, fullRows x 4: RGB residual + blend logit\nnetwork.dispose();\n```\n\nSee the [package README](https://github.com/bhouston/three-dlss-nr/blob/main/packages/three-dlss-nr/README.md) for the full API.\n\nThe network runs behind one interface, `NRBackend`, with two implementations that read the same features and write\nthe same head, so they can be switched at runtime and compared on identical inputs:\n\n| Factory | Entry point | What runs | Needs | \n|---|---|---|---|\n| `tslBackend` | `three-dlss-nr` | The native port: three.js TSL compute nodes, one `renderer.compute()` per frame | WebGPU, 24 KiB workgroup storage | \n| `referenceWgslBackend` | `three-dlss-nr/reference-backend` | **Upstream's own WebGPU port by maan, unchanged** (its JS and WGSL at`9d08f41` ) on three's device | `shader-f16` , 32 KiB workgroup storage | \n\nThe reference backend exists to compare speed and output against the original on the same renderer. Its code is\nOpenDLSS-NR's (MIT, Copyright (c) 2026 maan), copied byte for byte at build time with upstream's LICENSE and a\nSHA-256 provenance list. It is a separate entry point, so the main bundle does not carry it. Each factory's\n`unavailableReason(renderer)` returns `null` or a sentence naming what is missing and how to fix it.\n\nA third entry point, `three-dlss-nr/synthetic`, generates deterministic synthetic weights in the model directory\nlayout (`generateSyntheticModel`) and synthetic input features (` syntheticFeatures`), for tests and demos.\n\n**You supply the weights. None are included, downloaded, hosted or extracted.** The trained weights are NVIDIA's\nproprietary DLSS-NR 310.8.0 model. OpenDLSS-NR deliberately ships neither the weights nor instructions for obtaining\nthem, and grants no rights under NVIDIA's intellectual property. This project cannot redistribute, host or extract\nthem either: pulling them out of NVIDIA's binaries would likely breach NVIDIA's license terms. So neither this\nrepository, the npm package nor the demo contains them.\n\n**Synthetic weights** (`three-dlss-nr/synthetic`) are generated deterministically, in the same model directory layout,\nand calibrated so that every block boundary stays in a realistic, non-saturating range. They serve two purposes:\n\n- **They prove the port computes the same function as the reference.** The port is bit-exact against the reference\non arbitrary weights, through every kernel and every block, so it will match on any weights, including the real\nones.\n- **They let the whole pipeline run end to end** in the tests and in the demo.\n\nTheir output image is meaningless (the red field in the screenshot above).\n\n**If you are entitled to real weights,** point the network at the model directory: `manifest.json` plus\n`model/stages/*`, in the layout described in upstream's\n[`docs/weights.md`](https://github.com/maanHimself/OpenDLSS-NR/blob/9d08f41/docs/weights.md). In the demo, use\n**Load model directory...**; the browser reads it locally and never uploads it. In code, pass its URL or files to\n`NRModel.load`. The real-capture parity test (` network.real.gpu.test.ts`) runs upstream's own fixture comparison when\n`NR_WEIGHTS` (a model directory) and `NR_FIXTURES` (a fixture directory) point at local files.\n\nThere are two paths to meaningful public output: a license from NVIDIA, or open weights trained for the same\narchitecture. See [docs/training-weights.md](https://github.com/bhouston/three-dlss-nr/blob/main/docs/training-weights.md).\n\n**[three-dlss-nr.ben3d.ca](https://three-dlss-nr.ben3d.ca)** runs the network live in the browser on the\nLee Perry-Smith head scan. It needs a WebGPU browser (current Chrome or Edge, or Safari 26+).\n\n- **View:** NR off, NR on, or a split view with a draggable divider.\n- **Weights:** load your own model directory (read by the browser, never uploaded) or generate synthetic weights.\n- **Backend:** TSL (native three.js) or Reference WGSL (OpenDLSS-NR), with the network's GPU time per frame.\n- **NR settings:** intensity, style (none, cinematic, natural), local tone and structure, skin structure, auto mask,\ncolour strength and paper white, as in upstream's demo.\n- **Scene and device:** resolution up to 1280x720, and a report of the adapter's features and limits.\n\nThe site is [`packages/website`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/website) (TanStack Start, React, Tailwind), deployed to Cloud Run from\n`main`. Run it locally with `pnpm dev`.\n\n**Not real-time today.** The whole network takes hundreds of milliseconds per frame on an RTX 3060 Ti. The native\nTSL port is currently about 1.6x slower than the reference's hand-tuned WGSL on the same device, with the same bytes\nout. Closing that gap is future work.\n\nTSL vs reference WGSL, whole network, ms per frame: median GPU time (minimum in parentheses). RTX 3060 Ti, Chrome stable (D3D12 + DXC), synthetic weights, 451 dispatches per frame, 30 frames after 5 warm-up frames, timed with timestamp queries around the frame (wall time is within about 1.5 ms of GPU time).\n\n| Backend | 512x512 (field 576x512) | 1280x720 (field 1344x768) | \n|---|---|---|\n| TSL ( `tslBackend` ) | 201.2 (199.5) | 658.6 (655.2) | \n| Reference WGSL ( `referenceWgslBackend` ) | 123.4 (122.6) | 412.2 (410.8) | \n\n|  | TSL | Reference WGSL | \n|---|---|---|\n| Network creation (pipeline compile) | about 35 s | about 17 s | \n| First frame, 512x512 / 1280x720 | 420 / 845 ms | 165 / 648 ms | \n| Weights on the GPU | about 141 MiB | about 144 MiB | \n| Activations, 512x512 / 1280x720 | 545 / 1905 MiB | 545 / 1905 MiB | \n\nFor context, upstream reports 72-73 ms at 512x512 for its WebGPU port on an RTX 4070 SUPER, and 2.8 ms at 768x768 and 7.8 ms at 1920x1080 for its Vulkan FP8 tensor-core path on the same card. Those are upstream's numbers, measured on different hardware.\n\nReproduce on an otherwise idle GPU (the script runs both backends on one renderer in real Chrome, feeds them the same features and times them the same way):\n\n```\npnpm build && node scripts/bench-backends.mjs --sizes 512x512,1280x720\n```\n\n- three.js r180 or later (`three` is a peer dependency; developed and tested on 0.186) and its`WebGPURenderer` . There\nis no WebGL fallback.\n- A WebGPU device with at least 24 KiB of workgroup storage, 256 invocations per workgroup and 8 storage buffers per\nshader stage. `createNRDevice()` and`createNRRenderer()` request the adapter's maximum limits plus`shader-f16` and`timestamp-query` when available.\n- The reference backend also needs `shader-f16` and 32 KiB of workgroup storage.\n\n| Path | Contents | \n|---|---|\n| [`packages/three-dlss-nr`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/three-dlss-nr) | The library, published to npm. See its [README](https://github.com/bhouston/three-dlss-nr/blob/main/packages/three-dlss-nr/README.md) . | \n| [`packages/three-dlss-nr/src/README-internals.md`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/three-dlss-nr/src/README-internals.md) | How kernels are written in TSL and tested against the reference, and the TSL pitfalls. | \n| [`packages/website`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/website) | The demo site, deployed to Cloud Run. | \n| [`packages/fidelity-suite`](https://github.com/bhouston/three-dlss-nr/blob/main/packages/fidelity-suite) | The side-by-side parity suite (fidelity-kit) and its committed results, served at `/parity/` . | \n| [`reference/OpenDLSS-NR`](https://github.com/bhouston/three-dlss-nr/blob/main/reference/OpenDLSS-NR) | Upstream OpenDLSS-NR at `9d08f41` (git submodule, read-only), used for parity tests and bundled by the reference backend. | \n| [`docs/`](https://github.com/bhouston/three-dlss-nr/blob/main/docs) | Screenshots and [candidate head assets](https://github.com/bhouston/three-dlss-nr/blob/main/docs/suggested-assets.md) for the demo, with their licenses. | \n| [`scripts/`](https://github.com/bhouston/three-dlss-nr/blob/main/scripts) | Benchmark, bundle-size gate, release and CI helpers. | \n\n```\ngit clone --recurse-submodules https://github.com/bhouston/three-dlss-nr.git\ncd three-dlss-nr\ncorepack enable\npnpm install\npnpm build\npnpm dev           # library watch + demo site at http://localhost:3300\n```\n\n| Command | What it does | \n|---|---|\n| `pnpm build` | Bundle the reference port into `vendor/` , then build the library and the website | \n| `pnpm tsc` | Type-check the workspace | \n| `pnpm lint` | Oxlint | \n| `pnpm format` | Oxfmt ( `pnpm format:check` in CI) | \n| `pnpm test` | Type-check, then the unit tests (Node) with coverage | \n| `pnpm test:gpu` | The WebGPU tests in Node: the `gpu` project, then the whole network in the`gpu-network` project | \n| `pnpm size` | Minified gzip bundle-size gate, `three` external: 40 kB for`three-dlss-nr` , 8 kB for`three-dlss-nr/synthetic` | \n| `pnpm bench:backends` | Benchmark TSL against reference WGSL in real Chrome ( `--sizes 512x512,1280x720` ,`--model <dir>` ) | \n| `pnpm fidelity:check` | Check the committed parity results (every tensor and image bit-exact); `fidelity:generate` regenerates them in Chrome | \n| `pnpm fidelity:dev` | Browse the parity results locally with fidelity-kit | \n\n**GPU tests in Node.** `pnpm test:gpu` runs the `*.gpu.test.ts` files on headless WebGPU (Google's Dawn, via\n[`vitest-environment-webgpu-node`](https://github.com/bhouston/vitest-gpu)), with no browser. Dawn uses the platform\nbackend (D3D12 on Windows, Metal on macOS, Vulkan on Linux); set `DLSS_NR_DAWN_BACKEND=vulkan` or `DLSS_NR_DAWN` to\nchoose another. Linux without a GPU needs Mesa's software Vulkan driver, as in CI:\n\n```\nsudo apt-get install -y libegl1 libgles2 libgl1-mesa-dri mesa-vulkan-drivers\nexport LIBGL_ALWAYS_SOFTWARE=1\n```\n\nThe whole-network tests run at 64x64 by default. `NR_FULL=1` adds 512x512, whose single 451-dispatch command buffer\ncan trip the Windows driver timeout (TDR) when the GPU is busy. On Windows, Dawn's Vulkan backend loads only when\n`vulkan-1.dll` sits beside `node.exe` or `dawn.node`.\n\n**Parity in real Chrome.** Dawn in Node has no `shader-f16` on Windows, so reference parity runs in Chrome (stable,\nelse Playwright's Chromium) through a small DevTools harness:\n\n```\nnode packages/three-dlss-nr/test/browser/run-network-parity-chrome.mjs --sizes 64x64,512x512   # whole network, --golden re-pins CI digests\nnode packages/three-dlss-nr/test/browser/run-shim-parity-chrome.mjs                            # reference backend vs standalone upstream\nnode packages/three-dlss-nr/test/browser/run-fp8-gemm-chrome.mjs                               # FP8 GEMM against the CPU oracle\npnpm bench:backends                                                                            # TSL vs reference WGSL timings\npnpm fidelity:generate                                                                         # the side-by-side parity results (packages/fidelity-suite)\n```\n\nSee [CONTRIBUTING.md](https://github.com/bhouston/three-dlss-nr/blob/main/CONTRIBUTING.md) for the issue, branch, PR and commit workflow, the full list of local checks and\nthe release process, and [SECURITY.md](https://github.com/bhouston/three-dlss-nr/blob/main/SECURITY.md) for private vulnerability reports. Treat `reference/OpenDLSS-NR`\nas read-only, and never download, extract or commit NVIDIA model weights.\n\nMIT. [LICENSE](https://github.com/bhouston/three-dlss-nr/blob/main/LICENSE) carries both this project's copyright (Ben Houston and contributors, 2026) and OpenDLSS-NR's\nMIT notice (Copyright (c) 2026 maan). [NOTICE](https://github.com/bhouston/three-dlss-nr/blob/main/NOTICE) carries forward upstream's notice and lists the redistributed\nreference code and demo assets.\n\nThis project is not affiliated with, endorsed by, or supported by NVIDIA Corporation or by the author of OpenDLSS-NR. \"DLSS\" is a trademark of NVIDIA Corporation; it is used here only to describe what the network implemented by this code is compatible with. No NVIDIA software, weights, headers, or documentation are included, and no rights under any NVIDIA intellectual property are granted. You are responsible for the licenses that apply to whatever model data you use with it.\n\n- [OpenDLSS-NR](https://github.com/maanHimself/OpenDLSS-NR) by[maan](https://github.com/maanHimself) (MIT): the\nreimplementation of the network, its numerics, its documentation, and the WebGPU port that is this project's spec\nand its reference backend.\n- \"Infinite, 3D Head Scan\" by Lee Perry-Smith ([Infinite-Realities](https://ir-ltd.net) ), licensed under[CC BY 3.0](https://creativecommons.org/licenses/by/3.0/) : the demo's head. glTF and maps via the[three.js examples](https://github.com/mrdoob/three.js/tree/dev/examples/models/gltf/LeePerrySmith) .\n- [three.js](https://threejs.org) and its TSL node system, which the port is written in.\n\nCreated by [Ben Houston](https://ben3d.ca) and sponsored by [Land of Assets](https://landofassets.com).", "url": "https://wpnews.pro/news/threejs-port-of-opendlss-nr", "canonical_source": "https://github.com/bhouston/three-dlss-nr/", "published_at": "2026-10-03 02:15:22+00:00", "updated_at": "2026-10-03 02:36:07.618324+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "neural-networks", "generative-ai", "ai-research"], "entities": ["Three.js", "OpenDLSS-NR", "maanHimself", "Ben Houston", "NVIDIA", "DLSS 5", "WebGPU", "RTX 3060 Ti"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/threejs-port-of-opendlss-nr", "markdown": "https://wpnews.pro/news/threejs-port-of-opendlss-nr.md", "text": "https://wpnews.pro/news/threejs-port-of-opendlss-nr.txt", "jsonld": "https://wpnews.pro/news/threejs-port-of-opendlss-nr.jsonld"}}