{"slug": "show-hn-rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu", "title": "Show HN: Rgpu – a PyTorch device whose tensors live on a remote GPU", "summary": "A developer released rGPU, an open-source project that lets PyTorch programs run GPU work on a remote NVIDIA machine via a custom `rgpu` device or a CUDA shim covering libcuda, CUDA Runtime, cuBLAS, cuBLASLt and cuDNN. The project installs with `pip install rgpu` and is launched through `rgpu-run --host user@gpu-host --ssh-port 2222`, with a smoke test of `torch.ones(4, device=\"rgpu\")` expected to print 8.0. rGPU's documentation warns that neither protocol authenticates or encrypts connections, so the CUDA server's port 9713 must be restricted by firewall rules, and it is licensed under Apache License 2.0.", "body_md": "rGPU runs GPU work on a remote NVIDIA machine while the application stays on the client. It currently offers two paths:\n\n| Path | Use it for | Interface | \n|---|---|---|\n| PyTorch device | PyTorch programs that can opt into an `rgpu` device | `torch` operations over TCP | \n| CUDA shim | Existing Linux CUDA programs, including stock CUDA PyTorch | `libcuda` , CUDA Runtime, cuBLAS, cuBLASLt, and cuDNN shims | \n\nThe PyTorch device is the simpler integration. The CUDA shim covers existing binaries but has a larger compatibility surface.\n\nThe Fumadocs site in [`website/`](https://github.com/ymcrcat/rgpu/blob/main/website) is the product documentation:\n\nEngineering records and experiments are indexed in [`docs/README.md`](https://github.com/ymcrcat/rgpu/blob/main/docs/README.md).\n\nInstall rGPU with `pip install rgpu`, or `pip install -e ./python` from this\ncheckout, then follow the [quickstart](https://github.com/ymcrcat/rgpu/blob/main/website/content/docs/quickstart.mdx) to\ndeploy the server. Save this as `smoke.py` in your workload directory:\n\n``` python\nimport torch\nimport rgpu\n\nx = torch.ones(4, device=\"rgpu\")\nprint((x * 2).sum().item())  # 8.0\n```\n\nRun it in the environment where rGPU is installed, using your server's SSH destination and options:\n\n```\nrgpu-run --host user@gpu-host --ssh-port 2222 -i ~/.ssh/gpu_key \\\n  python smoke.py\n```\n\nThe program selects the device; `rgpu-run` opens the tunnel and configures the\nconnection. The expected output is `8.0`.\n\nFor existing Linux CUDA programs, follow the\n[CUDA shim guide](https://github.com/ymcrcat/rgpu/blob/main/website/content/docs/cuda-shim.mdx), starting with\n`./scripts/build_client.sh`.\n\nNeither protocol authenticates or encrypts connections. Keep `rgpu-opserver`\non its default localhost bind and use SSH. The CUDA server listens on all IPv4\ninterfaces: restrict port 9713 with host/cloud firewall rules before starting\nit, even when using an SSH tunnel. See [deployment](https://github.com/ymcrcat/rgpu/blob/main/website/content/docs/operations.mdx).\n\n```\n# C++ client and fake-driver tests\n./scripts/build_client.sh\n\n# Python tests\npython -m pip install -e './python[test]'\npython -m pytest python/tests\n\n# Static documentation\nnpm --prefix website ci\nnpm --prefix website run build\n```\n\nSee [`scripts/README.md`](https://github.com/ymcrcat/rgpu/blob/main/scripts/README.md) for the remaining build, cloud,\nand hardware commands. Generated C++ is committed; its policy and regeneration\nsteps are in [`codegen/README.md`](https://github.com/ymcrcat/rgpu/blob/main/codegen/README.md).\n\n| Path | Purpose | \n|---|---|\n| `client/` | CUDA client shims and transport | \n| `server/` | CUDA server and dispatch | \n| `common/` | Shared protocol and generated API metadata | \n| `python/` | PyTorch device and launcher | \n| `tests/` | C++, Python, CUDA, and hardware checks | \n| `codegen/` | CUDA header parser and source generators | \n| `website/` | Fumadocs product documentation | \n| `docs/` | Design records, measurements, and experiment reports | \n| `jax/` | Experimental JAX work; not a supported product path | \n| `scripts/` | Build, deployment, cloud, and test helpers | \n| `skills/` | Installable agent guidance for using rGPU | \n\nHistorical implementation notes and experimental results are indexed in\n[`docs/README.md`](https://github.com/ymcrcat/rgpu/blob/main/docs/README.md).\n\nLicensed under the [Apache License 2.0](https://github.com/ymcrcat/rgpu/blob/main/LICENSE).", "url": "https://wpnews.pro/news/show-hn-rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu", "canonical_source": "https://github.com/ymcrcat/rgpu", "published_at": "2026-10-07 05:15:44+00:00", "updated_at": "2026-10-07 05:49:51.873305+00:00", "lang": "en", "topics": ["ai-infrastructure", "developer-tools", "ai-tools"], "entities": ["rGPU", "PyTorch", "NVIDIA", "CUDA", "cuBLAS", "cuDNN", "JAX", "Apache License 2.0"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu", "markdown": "https://wpnews.pro/news/show-hn-rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu.md", "text": "https://wpnews.pro/news/show-hn-rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu.txt", "jsonld": "https://wpnews.pro/news/show-hn-rgpu-a-pytorch-device-whose-tensors-live-on-a-remote-gpu.jsonld"}}