Show HN: PHP-GPU-tensors – Native GPU operations in PHP Developer lcmialichi released PHP-GPU-tensors 0.1.0-beta.4, a native PHP extension that runs GPU tensor operations and runtime-compiled CUDA kernels without a Python runtime. In a PatchCamelyon MLP training workload on an NVIDIA GeForce MX570 A, the extension's optional kernel fusion ran 0.66 ms per step versus 4.08 ms eager, cutting 1,280 steps from 5.22 s to 0.84 s with identical metrics (accuracy 77.29%, ROC AUC 0.853). The library requires Linux, PHP 8.1–8.5, and an NVIDIA GPU, and offers no automatic differentiation or CPU fallback. GPU tensors, runtime-compiled CUDA kernels and fused tensor expressions for PHP. No Python runtime required. Native PHP extension for GPU tensors and NVIDIA CUDA-accelerated numerical workloads. Build tensor operations and machine-learning data pipelines in PHP, move data explicitly between host and GPU, and compile custom CUDA C++ kernels at runtime with NVRTC. Optionally fuse tensor expressions into compiled plans that replay quickly, with asynchronous execution and CUDA Graph for compatible plans. Status: beta 0.1.0-beta.4 · Linux · PHP 8.1–8.5, NTS and ZTS · NVIDIA GPU required. See Validation status validation-status for exactly what has been tested on real GPUs. - GPU tensors. CudaArray supports arithmetic, broadcasting, comparisons, matmul including batches , reductions, views and Python-style slice . - Explicit data movement. Packed buffers, .npy import and optional pinned host memory keep transfers cheap and visible. - Custom CUDA kernels. Compile CUDA C++ at runtime with NVRTC and launch it from PHP, synchronously or asynchronously. - Optional kernel fusion. Capture a PHP closure once, then replay it as fused GPU kernels. In a small training workload this was about 6× faster than eager execution with identical results see Results results . - Clear scope. A low-level GPU computing library, not a machine-learning framework: no automatic differentiation, no CPU fallback. Contents: Quick look quick-look · Results results · Install install · Tensors gpu-tensors-in-php · Slicing python-style-slicing · Data pipelines data-pipelines-for-machine-learning · Custom kernels custom-kernels · Fusion optional-kernel-fusion · Streams and CUDA Graph streams-asynchronous-execution-and-cuda-graph · Training example real-training-with-fusion · API and limits api-and-limits · Validation status validation-status · Contribute contribute php use Cuda\CudaArray; use Cuda\Fusion; $a = CudaArray::ones 4 ; $b = CudaArray::full 4 , 2.0 ; $c = CudaArray::full 4 , 3.0 ; // Eager execution default : one GPU operation per call. $eager = $a + $b $c; // Fusion opt-in : capture once, replay as fused GPU kernels. $plan = Fusion::compile fn $a, $b, $c = $a + $b $c, inputs: $a, $b, $c ; $fused = $plan- run $a, $b, $c ; print r $fused- toArray ; // 7, 7, 7, 7 Try a full training run on the GPU, written entirely in PHP: php -n -d extension=./cuda build-8.3/modules/cuda.so examples/08 gpu classifier.php \ --epochs=200 --batch-size=256 --learning-rate=0.05 --no-save Measured on PHP 8.3 NTS with an NVIDIA GeForce MX570 A 4 GB, compute capability 8.6 , driver 12.6 and CUDA runtime 12.3. The models are small, so these numbers mostly show how much per-step overhead Fusion removes; they are not peak GPU throughput and not a general-purpose GPU benchmark. Fusion vs. eager execution on the same model and data, with identical metrics accuracy 77.29%, ROC AUC 0.853, same confusion matrix in both modes : | Mode | Time per step | Patches per second | Training time 1,280 steps | |---|---|---|---| | Eager | 4.08 ms | 125,444 | 5.22 s | | Fusion compiled replay | 0.66 ms | 776,251 | 0.84 s | Workload: a small MLP hidden size 256 on 32,768 training and 4,096 test patches of PatchCamelyon https://github.com/basveeling/pcam CC0 , using 480 handcrafted features per patch RGB mean/std over an 8×8 grid and its central 4×4 , batch size 512, 20 epochs, learning rate 0.02. This is a performance demonstration, not a clinical model. The planner turned the 68 captured nodes of the training step into 11 fused kernels plus 9 native boundaries matmul and reductions . End-to-end training example examples/08 gpu classifier.php https://github.com/lcmialichi/php-gpu-tensors/blob/main/examples/08 gpu classifier.php , a complete Multi-Layer Perceptron MLP for MNIST, Fashion-MNIST, or custom CSVs. It processes hundreds of thousands of samples per second using fused CUDA kernels, showcasing tensor operations, JIT compilation, math optimization AdamW/SGD , and pure PHP data orchestration. Full benchmark reports live in the benchmarks repository benchmarks . | Need | Details | |---|---| | GPU and driver | CUDA-capable NVIDIA GPU; the host driver provides libcuda.so.1 at runtime | | Build toolchain | CUDA Toolkit including NVRTC , C/C++ toolchain, make , autoconf | | PHP | 8.1–8.5 development headers phpize , php-config , NTS or ZTS | | OS | Linux | Building requires the toolkit; running requires the host driver. git clone https://github.com/lcmialichi/php-gpu-tensors.git cd php-gpu-tensors ./compile.sh ./run-tests.sh --require-gpu php -n -d extension=./cuda build-8.3/modules/cuda.so examples/01 basics cuda array.php The build stays in cuda build-