cd /news/artificial-intelligence/threejs-port-of-opendlss-nr · home › topics › artificial-intelligence › article
[ARTICLE · art-144242] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

ThreeJS Port of OpenDLSS-NR

Developer Ben Houston released three-dlss-nr, a Three.js port of maanHimself's OpenDLSS-NR that runs the 71-block DLSS 5 neural rendering network as TSL compute kernels inside a WebGPURenderer with no CPU readback. The port tracks upstream commit 9d08f41 and is byte-identical to the reference WebGPU implementation across all 83 tensors and 451 compute dispatches at 64x64 and 512x512, verified in Chrome on an RTX 3060 Ti. It ships with synthetic weights only, so the demo output is not meaningful, and it is not affiliated with NVIDIA and includes no NVIDIA weights.

read13 min views1 publishedOct 3, 2026
ThreeJS Port of OpenDLSS-NR
Image: Michielbdejong (auto-discovered)

A port to Three.js (TSL / WebGPU) of OpenDLSS-NR by maan.

OpenDLSS-NR is an open-source reimplementation of the network behind NVIDIA's DLSS 5 Neural Rendering. All credit for that reimplementation, its numerics and its documentation belongs to OpenDLSS-NR. This repository ports it to three.js and tracks upstream commit 9d08f41, included as the reference/OpenDLSS-NR submodule. Not affiliated with NVIDIA, and no NVIDIA weights are included (details).

The DLSS 5 neural rendering network, as three.js compute nodes, byte for byte. three-dlss-nr runs the 71-block OpenDLSS-NR network as TSL compute kernels inside a WebGPURenderer, fed straight from your three.js scene with no CPU readback. Every intermediate tensor is byte-identical to the reference WebGPU port.

Synthetic weights: output is not meaningful. The demo in split view on the TSL backend, NR off on the left and NR on on the right. Synthetic weights have the right layout and run the network's real arithmetic, but the image they produce means nothing, hence the red field (why). With a real model directory the right half is the re-rendered head. Head: Lee Perry-Smith, CC BY 3.0 (credits).

Try it live at three-dlss-nr.ben3d.ca.

From upstream's README: it is a generative neural rendering network (NVIDIA's term). It re-renders the frame the engine already drew, generating detail from injected noise and adjusting tone, structure and skin under a style setting. Input and output are the same resolution; it is not an upscaler.

  • A U-net of shifted-window (Swin) transformer blocks with a global ViT at the bottom: 71 blocks over six pooling levels, FP8 (E4M3) activations with FP16 accumulation, 141 MiB of weights.
  • In: one rendered frame (a low dynamic range proxy of it, three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars). Out: four f32 channels per pixel, an RGB residual and one temporal-blend logit.
  • 451 compute dispatches per frame in the WebGPU formulation.

NVIDIA describes the model in its report, DLSS 5: Generative Neural Rendering (project page). Upstream's docs/network.md is the graph in full.

The spec for every kernel is upstream's WebGPU port (ports/browser-webgpu), and the port is checked against it on the same device, the same inputs and the same weights:

  • Every kernel family (FP8 and f16 GEMMs, window attention, the global ViT, the elementwise ops, preprocess, and the frame kernels that build the input features and compose the output) is byte-identical to the reference.
  • The whole network (451 dispatches) is byte-identical at 64x64 and at 512x512 on synthetic weights: all 79 block boundaries, the three post tensors and the head, 83 of 83 tensors, with repeat runs identical. Checked in real Chrome on an RTX 3060 Ti (D3D12 + DXC), with both networks in one page on one device.
  • The numerics (E4M3 encode and decode, f16 rounding, the SiLU and exponent tables, FP8 and f16 dot products) are checked on the GPU against all seven sections of the reference's numerics fixture, which upstream's Vulkan implementation produced: exhaustively over every half and every byte where the domain allows it.
  • CI (Linux, Mesa lavapipe) runs the whole network in Node and gates it against golden digests recorded from the Chrome run. It cannot compare against the reference directly: llvmpipe folds thef32(f16(x)) round trip intox , so the reference's f16 roundings silently do not happen there. The TSL port never relies on f16 hardware. It rounds on f32 bit patterns, so it produces the same bytes on lavapipe as on a real GPU.

See it side by side: three-dlss-nr.ben3d.ca/parity/ shows the reference running standalone, the TSL port and the reference shim on the same frames (the head scan and two procedural scenes, at 256x256 and 512x512): the composed output, the head and four block boundaries, all byte-identical, with every raw tensor compared and per-scene GPU times. It is built with fidelity-kit from packages/fidelity-suite, and CI checks its committed verdicts.

A consequence: the TSL port does not need shader-f16, while the reference does. How the kernels are written and tested is in src/README-internals.md.

npm install three-dlss-nr three

DlssNrPass puts the network behind a three.js scene. It renders the scene into an RGBA16F target (linear colour plus three's velocity), builds the network's input features on the GPU, runs the network, composes the result with the temporal history, and draws NR off, NR on or a split view onto the canvas.

import { backendBuilder, createNRRenderer, DlssNrPass, NRModel, tslBackend } from 'three-dlss-nr';

const { renderer } = await createNRRenderer(); // a WebGPURenderer on a device with the limits the network needs
document.body.appendChild(renderer.domElement);

// Your model directory: manifest.json + model/stages/* (see the weights section below).
const model = await NRModel.load('https://example.com/my-model');

const pass = new DlssNrPass({ renderer, scene, camera, width: 960, height: 540, view: 'split' });
await pass.setNetwork(backendBuilder(tslBackend, model)); // compiles the kernels; setSize() rebuilds
pass.setSettings({ intensity: 1, style: 1 }); // style: 0 none, 1 cinematic, 2 natural

renderer.setAnimationLoop(async () => {
  const stats = await pass.render(); // null while the previous frame is still on the GPU
  // stats?.network: { method, gpuMilliseconds, wallMilliseconds }
});

Load the model once and share it: resizes and network rebuilds reuse it instead of re-reading 141 MiB. Call pass.resetHistory() on a camera cut. The internal size is capped at 1280x720.

To run the network on your own inputs instead of a scene, use a backend directly:

const network = await tslBackend.create({ renderer, model, width: 512, height: 512, onProgress: console.log });
network.writeFeatures(features); // Float32Array, fullRows x 16 lanes (or write network.features from a compute pass)
const timing = await network.run({ timing: true });
const head = await network.readHead(); // Float32Array, fullRows x 4: RGB residual + blend logit
network.dispose();

See the package README for the full API.

The network runs behind one interface, NRBackend, with two implementations that read the same features and write the same head, so they can be switched at runtime and compared on identical inputs:

Factory Entry point What runs Needs
tslBackend three-dlss-nr The native port: three.js TSL compute nodes, one renderer.compute() per frame WebGPU, 24 KiB workgroup storage
referenceWgslBackend three-dlss-nr/reference-backend Upstream's own WebGPU port by maan, unchanged (its JS and WGSL at9d08f41 ) on three's device shader-f16 , 32 KiB workgroup storage

The reference backend exists to compare speed and output against the original on the same renderer. Its code is OpenDLSS-NR's (MIT, Copyright (c) 2026 maan), copied byte for byte at build time with upstream's LICENSE and a SHA-256 provenance list. It is a separate entry point, so the main bundle does not carry it. Each factory's unavailableReason(renderer) returns null or a sentence naming what is missing and how to fix it.

A third entry point, three-dlss-nr/synthetic, generates deterministic synthetic weights in the model directory layout (generateSyntheticModel) and synthetic input features ( syntheticFeatures), for tests and demos.

You supply the weights. None are included, downloaded, hosted or extracted. The trained weights are NVIDIA's proprietary DLSS-NR 310.8.0 model. OpenDLSS-NR deliberately ships neither the weights nor instructions for obtaining them, and grants no rights under NVIDIA's intellectual property. This project cannot redistribute, host or extract them either: pulling them out of NVIDIA's binaries would likely breach NVIDIA's license terms. So neither this repository, the npm package nor the demo contains them.

Synthetic weights (three-dlss-nr/synthetic) are generated deterministically, in the same model directory layout, and calibrated so that every block boundary stays in a realistic, non-saturating range. They serve two purposes:

  • They prove the port computes the same function as the reference. The port is bit-exact against the reference on arbitrary weights, through every kernel and every block, so it will match on any weights, including the real ones.
  • They let the whole pipeline run end to end in the tests and in the demo.

Their output image is meaningless (the red field in the screenshot above).

If you are entitled to real weights, point the network at the model directory: manifest.json plus model/stages/*, in the layout described in upstream's docs/weights.md. In the demo, use Load model directory...; the browser reads it locally and never uploads it. In code, pass its URL or files to NRModel.load. The real-capture parity test ( network.real.gpu.test.ts) runs upstream's own fixture comparison when NR_WEIGHTS (a model directory) and NR_FIXTURES (a fixture directory) point at local files.

There are two paths to meaningful public output: a license from NVIDIA, or open weights trained for the same architecture. See docs/training-weights.md.

three-dlss-nr.ben3d.ca runs the network live in the browser on the Lee Perry-Smith head scan. It needs a WebGPU browser (current Chrome or Edge, or Safari 26+).

  • View: NR off, NR on, or a split view with a draggable divider.
  • Weights: load your own model directory (read by the browser, never uploaded) or generate synthetic weights.
  • Backend: TSL (native three.js) or Reference WGSL (OpenDLSS-NR), with the network's GPU time per frame.
  • NR settings: intensity, style (none, cinematic, natural), local tone and structure, skin structure, auto mask, colour strength and paper white, as in upstream's demo.
  • Scene and device: resolution up to 1280x720, and a report of the adapter's features and limits.

The site is packages/website (TanStack Start, React, Tailwind), deployed to Cloud Run from main. Run it locally with pnpm dev.

Not real-time today. The whole network takes hundreds of milliseconds per frame on an RTX 3060 Ti. The native TSL port is currently about 1.6x slower than the reference's hand-tuned WGSL on the same device, with the same bytes out. Closing that gap is future work.

TSL vs reference WGSL, whole network, ms per frame: median GPU time (minimum in parentheses). RTX 3060 Ti, Chrome stable (D3D12 + DXC), synthetic weights, 451 dispatches per frame, 30 frames after 5 warm-up frames, timed with timestamp queries around the frame (wall time is within about 1.5 ms of GPU time).

Backend 512x512 (field 576x512) 1280x720 (field 1344x768)
TSL ( tslBackend ) 201.2 (199.5) 658.6 (655.2)
Reference WGSL ( referenceWgslBackend ) 123.4 (122.6) 412.2 (410.8)
TSL Reference WGSL
Network creation (pipeline compile) about 35 s about 17 s
First frame, 512x512 / 1280x720 420 / 845 ms 165 / 648 ms
Weights on the GPU about 141 MiB about 144 MiB
Activations, 512x512 / 1280x720 545 / 1905 MiB 545 / 1905 MiB

For context, upstream reports 72-73 ms at 512x512 for its WebGPU port on an RTX 4070 SUPER, and 2.8 ms at 768x768 and 7.8 ms at 1920x1080 for its Vulkan FP8 tensor-core path on the same card. Those are upstream's numbers, measured on different hardware.

Reproduce on an otherwise idle GPU (the script runs both backends on one renderer in real Chrome, feeds them the same features and times them the same way):

pnpm build && node scripts/bench-backends.mjs --sizes 512x512,1280x720
  • three.js r180 or later (three is a peer dependency; developed and tested on 0.186) and itsWebGPURenderer . There is no WebGL fallback.
  • A WebGPU device with at least 24 KiB of workgroup storage, 256 invocations per workgroup and 8 storage buffers per shader stage. createNRDevice() andcreateNRRenderer() request the adapter's maximum limits plusshader-f16 andtimestamp-query when available.
  • The reference backend also needs shader-f16 and 32 KiB of workgroup storage.
Path Contents
packages/three-dlss-nr The library, published to npm. See its README .
packages/three-dlss-nr/src/README-internals.md How kernels are written in TSL and tested against the reference, and the TSL pitfalls.
packages/website The demo site, deployed to Cloud Run.
packages/fidelity-suite The side-by-side parity suite (fidelity-kit) and its committed results, served at /parity/ .
reference/OpenDLSS-NR Upstream OpenDLSS-NR at 9d08f41 (git submodule, read-only), used for parity tests and bundled by the reference backend.
docs/ Screenshots and candidate head assets for the demo, with their licenses.
scripts/ Benchmark, bundle-size gate, release and CI helpers.
git clone --recurse-submodules https://github.com/bhouston/three-dlss-nr.git
cd three-dlss-nr
corepack enable
pnpm install
pnpm build
pnpm dev           # library watch + demo site at http://localhost:3300
Command What it does
pnpm build Bundle the reference port into vendor/ , then build the library and the website
pnpm tsc Type-check the workspace
pnpm lint Oxlint
pnpm format Oxfmt ( pnpm format:check in CI)
pnpm test Type-check, then the unit tests (Node) with coverage
pnpm test:gpu The WebGPU tests in Node: the gpu project, then the whole network in thegpu-network project
pnpm size Minified gzip bundle-size gate, three external: 40 kB forthree-dlss-nr , 8 kB forthree-dlss-nr/synthetic
pnpm bench:backends Benchmark TSL against reference WGSL in real Chrome ( --sizes 512x512,1280x720 ,--model <dir> )
pnpm fidelity:check Check the committed parity results (every tensor and image bit-exact); fidelity:generate regenerates them in Chrome
pnpm fidelity:dev Browse the parity results locally with fidelity-kit

GPU tests in Node. pnpm test:gpu runs the *.gpu.test.ts files on headless WebGPU (Google's Dawn, via vitest-environment-webgpu-node), with no browser. Dawn uses the platform backend (D3D12 on Windows, Metal on macOS, Vulkan on Linux); set DLSS_NR_DAWN_BACKEND=vulkan or DLSS_NR_DAWN to choose another. Linux without a GPU needs Mesa's software Vulkan driver, as in CI:

sudo apt-get install -y libegl1 libgles2 libgl1-mesa-dri mesa-vulkan-drivers
export LIBGL_ALWAYS_SOFTWARE=1

The whole-network tests run at 64x64 by default. NR_FULL=1 adds 512x512, whose single 451-dispatch command buffer can trip the Windows driver timeout (TDR) when the GPU is busy. On Windows, Dawn's Vulkan backend loads only when vulkan-1.dll sits beside node.exe or dawn.node.

Parity in real Chrome. Dawn in Node has no shader-f16 on Windows, so reference parity runs in Chrome (stable, else Playwright's Chromium) through a small DevTools harness:

node packages/three-dlss-nr/test/browser/run-network-parity-chrome.mjs --sizes 64x64,512x512   # whole network, --golden re-pins CI digests
node packages/three-dlss-nr/test/browser/run-shim-parity-chrome.mjs                            # reference backend vs standalone upstream
node packages/three-dlss-nr/test/browser/run-fp8-gemm-chrome.mjs                               # FP8 GEMM against the CPU oracle
pnpm bench:backends                                                                            # TSL vs reference WGSL timings
pnpm fidelity:generate                                                                         # the side-by-side parity results (packages/fidelity-suite)

See CONTRIBUTING.md for the issue, branch, PR and commit workflow, the full list of local checks and the release process, and SECURITY.md for private vulnerability reports. Treat reference/OpenDLSS-NR as read-only, and never download, extract or commit NVIDIA model weights.

MIT. LICENSE carries both this project's copyright (Ben Houston and contributors, 2026) and OpenDLSS-NR's MIT notice (Copyright (c) 2026 maan). NOTICE carries forward upstream's notice and lists the redistributed reference code and demo assets.

This project is not affiliated with, endorsed by, or supported by NVIDIA Corporation or by the author of OpenDLSS-NR. "DLSS" is a trademark of NVIDIA Corporation; it is used here only to describe what the network implemented by this code is compatible with. No NVIDIA software, weights, headers, or documentation are included, and no rights under any NVIDIA intellectual property are granted. You are responsible for the licenses that apply to whatever model data you use with it.

  • OpenDLSS-NR bymaan (MIT): the reimplementation of the network, its numerics, its documentation, and the WebGPU port that is this project's spec and its reference backend.
  • "Infinite, 3D Head Scan" by Lee Perry-Smith (Infinite-Realities ), licensed underCC BY 3.0 : the demo's head. glTF and maps via thethree.js examples .
  • three.js and its TSL node system, which the port is written in.

Created by Ben Houston and sponsored by Land of Assets.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @three.js 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/threejs-port-of-open…] indexed:0 read:13min 2026-10-03 · —