{"slug": "show-hn-php-gpu-tensors-native-gpu-operations-in-php", "title": "Show HN: PHP-GPU-tensors – Native GPU operations in PHP", "summary": "Developer lcmialichi released PHP-GPU-tensors 0.1.0-beta.4, a native PHP extension that runs GPU tensor operations and runtime-compiled CUDA kernels without a Python runtime. In a PatchCamelyon MLP training workload on an NVIDIA GeForce MX570 A, the extension's optional kernel fusion ran 0.66 ms per step versus 4.08 ms eager, cutting 1,280 steps from 5.22 s to 0.84 s with identical metrics (accuracy 77.29%, ROC AUC 0.853). The library requires Linux, PHP 8.1–8.5, and an NVIDIA GPU, and offers no automatic differentiation or CPU fallback.", "body_md": "**GPU tensors, runtime-compiled CUDA kernels and fused tensor expressions for PHP. No Python runtime required.**\n\nNative PHP extension for GPU tensors and NVIDIA CUDA-accelerated numerical workloads. Build tensor operations and machine-learning data pipelines in PHP, move data explicitly between host and GPU, and compile custom CUDA C++ kernels at runtime with NVRTC. Optionally fuse tensor expressions into compiled plans that replay quickly, with asynchronous execution and CUDA Graph for compatible plans.\n\n**Status:** beta (`0.1.0-beta.4`) · Linux · PHP 8.1–8.5, NTS and ZTS · NVIDIA GPU\nrequired. See [Validation status](#validation-status) for exactly what has been\ntested on real GPUs.\n\n- **GPU tensors.**`CudaArray` supports arithmetic, broadcasting, comparisons,`matmul()` (including batches), reductions, views and Python-style`slice()` .\n- **Explicit data movement.** Packed buffers,`.npy` import and optional pinned\nhost memory keep transfers cheap and visible.\n- **Custom CUDA kernels.** Compile CUDA C++ at runtime with NVRTC and launch it\nfrom PHP, synchronously or asynchronously.\n- **Optional kernel fusion.** Capture a PHP closure once, then replay it as fused\nGPU kernels. In a small training workload this was about 6× faster than eager\nexecution with identical results (see[Results](#results) ).\n- **Clear scope.** A low-level GPU computing library, not a machine-learning\nframework: no automatic differentiation, no CPU fallback.\n\n**Contents:**\n[Quick look](#quick-look) ·\n[Results](#results) ·\n[Install](#install) ·\n[Tensors](#gpu-tensors-in-php) ·\n[Slicing](#python-style-slicing) ·\n[Data pipelines](#data-pipelines-for-machine-learning) ·\n[Custom kernels](#custom-kernels) ·\n[Fusion](#optional-kernel-fusion) ·\n[Streams and CUDA Graph](#streams-asynchronous-execution-and-cuda-graph) ·\n[Training example](#real-training-with-fusion) ·\n[API and limits](#api-and-limits) ·\n[Validation status](#validation-status) ·\n[Contribute](#contribute)\n\n``` php\nuse Cuda\\CudaArray;\nuse Cuda\\Fusion;\n\n$a = CudaArray::ones([4]);\n$b = CudaArray::full([4], 2.0);\n$c = CudaArray::full([4], 3.0);\n\n// Eager execution (default): one GPU operation per call.\n$eager = $a + $b * $c;\n\n// Fusion (opt-in): capture once, replay as fused GPU kernels.\n$plan = Fusion::compile(fn($a, $b, $c) => $a + $b * $c, inputs: [$a, $b, $c]);\n$fused = $plan->run($a, $b, $c);\n\nprint_r($fused->toArray()); // [7, 7, 7, 7]\n```\n\nTry a full training run on the GPU, written entirely in PHP:\n\n```\nphp -n -d extension=./cuda_build-8.3/modules/cuda.so examples/08_gpu_classifier.php \\\n  --epochs=200 --batch-size=256 --learning-rate=0.05 --no-save\n```\n\nMeasured on PHP 8.3 NTS with an NVIDIA GeForce MX570 A (4 GB, compute capability 8.6), driver 12.6 and CUDA runtime 12.3. The models are small, so these numbers mostly show how much per-step overhead Fusion removes; they are not peak GPU throughput and not a general-purpose GPU benchmark.\n\n**Fusion vs. eager execution** on the same model and data, with identical\nmetrics (accuracy 77.29%, ROC AUC 0.853, same confusion matrix in both modes):\n\n| Mode | Time per step | Patches per second | Training time (1,280 steps) | \n|---|---|---|---|\n| Eager | 4.08 ms | 125,444 | 5.22 s | \n| Fusion (compiled replay) | 0.66 ms | 776,251 | 0.84 s | \n\nWorkload: a small MLP (hidden size 256) on 32,768 training and 4,096 test patches\nof [PatchCamelyon](https://github.com/basveeling/pcam) (CC0), using 480\nhandcrafted features per patch (RGB mean/std over an 8×8 grid and its central\n4×4), batch size 512, 20 epochs, learning rate 0.02. This is a performance\ndemonstration, not a clinical model. The planner turned the 68 captured nodes of\nthe training step into 11 fused kernels plus 9 native boundaries (matmul and\nreductions).\n\n**End-to-end training example** ([`examples/08_gpu_classifier.php`](https://github.com/lcmialichi/php-gpu-tensors/blob/main/examples/08_gpu_classifier.php), a complete Multi-Layer Perceptron (MLP) for MNIST, Fashion-MNIST, or custom CSVs. It processes hundreds of thousands of samples per second using fused CUDA kernels, showcasing tensor operations, JIT compilation, math optimization (AdamW/SGD), and pure PHP data orchestration.\n\nFull benchmark reports live in the\n[benchmarks repository](#benchmarks).\n\n| Need | Details | \n|---|---|\n| GPU and driver | CUDA-capable NVIDIA GPU; the host driver provides `libcuda.so.1` at runtime | \n| Build toolchain | CUDA Toolkit (including NVRTC), C/C++ toolchain, `make` ,`autoconf` | \n| PHP | 8.1–8.5 development headers ( `phpize` ,`php-config` ), NTS or ZTS | \n| OS | Linux | \n\nBuilding requires the toolkit; running requires the host driver.\n\n```\ngit clone https://github.com/lcmialichi/php-gpu-tensors.git\ncd php-gpu-tensors\n./compile.sh\n./run-tests.sh --require-gpu\nphp -n -d extension=./cuda_build-8.3/modules/cuda.so examples/01_basics_cuda_array.php\n```\n\nThe build stays in `cuda_build-<PHP major.minor>/modules/cuda.so`; replace `8.3`\nabove with the PHP version used by `php-config`. To install the extension and its\nINI configuration instead, run `./compile.sh --install` with permission to write\nto your PHP extension/INI directories. To choose another PHP ABI:\n\n```\nPHP_BIN=php8.3 PHPIZE=phpize8.3 PHP_CONFIG=php-config8.3 ./compile.sh\nPHP_BIN=php8.3 PHP_CONFIG=php-config8.3 ./run-tests.sh --require-gpu\n```\n\nBuild options:\n\n- `CUDA_HOME` if the toolkit is not at`/usr/local/cuda` .\n- `CUDA_ARCH=sm_86` when cross-building.\n- `CUDA_USE_CUBLAS=no ./compile.sh` to use the extension's built-in CUDA matrix\nkernels instead of cuBLAS (cuBLAS is used when available for compatible larger\nmatrix products).\n\nThe extension is published as\n[`lcmialichi/php-gpu-tensors`](https://packagist.org/packages/lcmialichi/php-gpu-tensors).\n\n```\npie install lcmialichi/php-gpu-tensors:0.1.0-beta.4\n```\n\nUse `0.1.0-beta.4` or newer for the Fusion APIs and the training example\ndescribed here; older releases may not include them. PIE builds the native\nextension for the selected PHP installation; it does not install an NVIDIA driver\nor CUDA Toolkit. Those must already be available on the system, and GPU execution\nadditionally requires a compatible NVIDIA driver and a visible GPU.\n\nUsers with the NVIDIA Container Toolkit and a working host driver can build and test in the development image:\n\n```\ndocker compose run --rm php_cuda_dev bash -lc './compile.sh && ./run-tests.sh --require-gpu'\n```\n\n`./run-tests.sh` runs CPU-side C tests and the PHP test suite. Use `--cpu-only` to\nrun only the host-side C tests in build environments without a GPU; this does not\nvalidate CUDA execution. `--require-gpu` fails immediately when no GPU is visible,\ninstead of treating skipped GPU tests as success.\n\n``` php\nuse Cuda\\CudaArray;\n\n$input = new CudaArray([[1, 2], [3, 4]], 'float32');\n$weights = CudaArray::ones([2, 2]);\n$output = $input->add($weights)->multiply(2);\n$column_means = $input->mean(0);\n\nprint_r($output->toArray()); // [[4, 6], [8, 10]]\necho $output->dtype();       // float32\n```\n\n`CudaArray` holds GPU storage; operations return GPU tensors. `toArray()`\ntransfers to the CPU and expands all values into PHP arrays. Operations include\narithmetic, broadcasting, comparisons, `matmul()` (including batches), shape\nviews, and `sum()`, `mean()`, `min()`, `max()`, `prod()`, `argMax()`, and\n`argMin()` with an optional axis. PHP arithmetic operators also dispatch to tensor\nmethods. `mean()` reduces all values when called without an axis, or reduces one\ndimension when given an axis. It returns `float32` for `float32` input and\n`float64` for `float64`, integer, and boolean input.\n\n`CudaArray::slice()` accepts a comma-separated expression, or one selector per\naxis. Ranges have an exclusive stop, omitted bounds default to the axis limits,\nand negative indices/bounds count from the end. Range bounds are clipped to the\naxis size; individual indices outside the axis throw `InvalidArgumentException`.\n\n``` php\n$batch = $tensor->slice('10:20, :');\n$columns = $tensor->slice(':, 2:8:2');\n$lastRows = $tensor->slice('-10:');\n$row = $tensor->slice(3);       // Removes the first axis.\n$oneRow = $tensor->slice('3:4'); // Keeps that axis with size 1.\n\n// Dynamic equivalents: [start, stop] or [start, stop, step].\n$batch = $tensor->slice([$start, $stop], null);\n$columns = $tensor->slice(null, [2, 8, 2]);\n```\n\nIntegers remove axes; `null`, `':'`, omitted axes and `slice()` with no arguments\nselect the full extent. Steps must be positive integers no larger than `INT_MAX`.\nZero/negative steps, ellipsis, new axes, index lists and executable expressions\nare not supported. Expressions are parsed as integers/ranges, never evaluated as\nPHP.\n\n## Views, empty ranges and legacy syntax\n\nA fully indexed tensor has shape `[]` and size 1; `toArray()` and\n`toHost()->toArray()` represent it as a one-element array.\n\nOutside Fusion, slices are zero-copy views with shared storage: writes through a view affect its parent, and the parent remains alive while the view is retained. Host transfers pack strided views into contiguous host storage. Row assignments involving strided tensors stage the source on the host before writing, preserving overlapping-source correctness; this is not a GPU-only bulk scatter operation. Fusion captures slice index transformations without materializing during capture; returned compiled/scoped outputs use the existing independent-output storage contract.\n\nEmpty ranges are supported, preserving the remaining shape: `[0, 4]` converts to\n`[]`, whereas `[3, 0]` converts to `[[], [], []]`. Elementwise operations, casts,\ntransfers, reshape/transpose and matmul handle zero elements. Sum/product over an\nempty axis return 0/1; mean returns NaN. Min/max and arg reductions throw when an\nempty reduced axis would produce values, because no identity/index exists. An\nempty non-reduced output remains empty. Packed-buffer imports accept zero\ndimensions; materialized empty results support serialization.\n\nThe legacy `__invoke()` and `[]` selection syntax remain unchanged, including\ntheir inclusive ranges. Do not interpret their bounds as the new exclusive\n`slice()` bounds.\n\nUse this extension to build GPU-accelerated numerical steps into PHP\napplications: tensor arithmetic, matrix multiplication, broadcasting, reductions\nsuch as `mean()`, and custom CUDA kernels. These primitives support\nmachine-learning data preparation, inference and small training loops while the\ndata remains in NVIDIA GPU memory.\n\nThis is a low-level GPU computing library, not a complete machine-learning framework. The training example implements backpropagation explicitly; automatic differentiation and Python interoperability are not provided.\n\nFor data already in packed row-major bytes, avoid creating individual PHP\nscalars. `fromFile()` reads raw bytes, whereas `fromNpy()` parses NumPy's `.npy`\nformat:\n\n``` php\nuse Cuda\\CudaArray;\nuse Cuda\\HostArray;\n\n$bytes = pack('g*', 1, 2, 3, 4); // little-endian float32\n$host = HostArray::fromBuffer($bytes, [2, 2]);\n$gpu = $host->toGpu();\n$result = CudaArray::where(\n    CudaArray::fromBuffer(pack('C*', 1, 0), [2], 'bool'),\n    $gpu,\n    CudaArray::zeros([2, 2])\n);\n$cpuCopy = $result->toHost();\n$raw = $cpuCopy->toBuffer();\n```\n\n`HostArray` is an alias of `Cuda\\ContiguousArray`, a contiguous CPU tensor. Pass\n`pinned: true` to its constructor or `fromBuffer()` for page-locked host storage\nwhen repeated transfers justify the extra host memory. `where()` broadcasts its\nthree inputs; its mask treats nonzero values as true, and its two value tensors\nmust have the same dtype. `.npy` imports support C-order little-endian numeric and\nboolean arrays; Fortran order and big-endian data are rejected.\n\nRegister the CUDA kernel's argument metadata, compile to PTX, then launch with explicit grid and block dimensions:\n\n``` php\nuse Cuda\\Compiler;\nuse Cuda\\CudaArray;\n\n$source = <<<'CUDA'\nextern \"C\" __global__ void scale(float *data, int factor, int count)\n{\n    int index = blockIdx.x * blockDim.x + threadIdx.x;\n    if (index < count) data[index] *= factor;\n}\nCUDA;\n\n$compiler = new Compiler(source: $source);\n$compiler->kernel('scale', [\n    ['name' => 'data', 'type' => 'array', 'dtype' => 'float32'],\n    ['name' => 'factor', 'dtype' => 'int32'],\n    ['name' => 'count', 'dtype' => 'int32'],\n]);\n$module = $compiler->compile();\n$module->initialize();\n\n$data = CudaArray::ones([512]);\n$module->launch('scale',\n    args: [$data, 3, 512],\n    config: ['block' => [256, 1, 1], 'grid' => [2, 1, 1]]\n);\nprint_r(array_slice($data->toArray(), 0, 4)); // [3, 3, 3, 3]\n```\n\n`launch()` synchronizes; `launchAsync()` returns an operation ID for `sync()` or\n`wait()`. Keep tensors alive until asynchronous work finishes. See\n[the JIT examples](https://github.com/lcmialichi/php-gpu-tensors/blob/main/examples/04_custom_jit_kernels.php) and\n[asynchronous execution](https://github.com/lcmialichi/php-gpu-tensors/blob/main/examples/05_jit_async_execution.php).\n\nEager execution remains the default. Fusion is explicitly opt-in: existing code keeps using eager execution unless capture is enabled.\n\n``` php\nflowchart LR\n    A[PHP closure] --> B[Capture once<br/>metadata-only placeholders]\n    B --> C[Plan]\n    C --> D[Fused elementwise kernels<br/>NVRTC to PTX]\n    C --> E[Native boundaries<br/>matmul and reductions]\n    D --> F[Replay on a private stream]\n    E --> F\n```\n\n`Fusion::run()` captures tensor operations inside a callback and returns\nmaterialized `CudaArray` outputs:\n\n``` php\nuse Cuda\\Fusion;\n\n$result = Fusion::run(fn() => $tensorA + $tensorB * $tensorC);\n$eager = Fusion::run(fn() => $tensorA + $tensorB * $tensorC, enabled: false);\n```\n\nFor repeated execution, compile a specialized plan once:\n\n``` php\n$graph = Fusion::compile(\n    fn($a, $b, $c) => $a + $b * $c,\n    inputs: [$tensorA, $tensorB, $tensorC]\n);\n$result = $graph->run($tensorA, $tensorB, $tensorC);\n$other = $graph->run($otherA, $otherB, $otherC);\nprint_r($graph->getStats());\nprint_r($graph->getPlan());\necho $graph->getSource();\n```\n\nWhat to know at a glance:\n\n- **Fused:** addition, subtraction, multiplication, division, unary operations,\ncomparisons,`where()` , safe`astype()` conversions and\nreshape/transpose/slice index transformations.\n- **Execution boundaries:** reductions,`matmul()` (currently float32) and powers\nrun on existing kernels; the fused kernels run in dependency order around them.\n- **Replay:** the PHP callback is not invoked again after compilation, and input\nvalues and pointers can change between executions.\n- **Inspection:**`getStats()` ,`getPlan()` and`getSource()` show what the planner\nproduced; for`$a + $b * $c` the plan has one fused kernel and no intermediate\ndata buffers.\n\n## Fusion reference: compile and replay, fusion rules, capture limits and caching\n\n**Compile and replay.** Compilation invokes the callback once with metadata-only\nplaceholders. Replay does not invoke PHP callback code again. Inputs must match the\nexample shapes, dtypes and strides, and execution must use the compilation device\nand CUDA context. Input values and pointers can change between executions; keep the\ndevice/context alive until the graph is released. Tensors captured by a closure are\nretained by the graph; use callback parameters for replaceable inputs.\n\n**What fuses.** Addition, subtraction, multiplication, division, unary operations\nand comparison methods fuse into elementwise kernels, with broadcasting,\nstrided/view inputs, scalar operands and dtype promotion. Each node converts to its\nown result dtype, preserving intermediate rounding/narrowing. Generated kernels use\nthe eager backend's fast-math settings but disable cross-node FMA contraction.\n`where()`, safe explicit `astype()` conversions and reshape/transpose/slice index\ntransformations also fuse. Safe casts work in eager execution too, using the same\ngenerated conversion kernel; unsafe narrowing retains the existing rejection\npolicy.\n\n**Boundaries and planning.** Reductions, `matmul()` (currently float32) and powers\nare execution boundaries using existing kernels. All generated elementwise kernels\nin a plan are compiled together through the existing `Compiler` NVRTC\ninfrastructure, then executed in dependency order around these boundaries.\nExpressions split at a weighted cost budget of 32 (math functions, index transforms\nand float64 have higher cost). Shared expensive expressions can be materialized\ninstead of recomputed. Pure repeated binary/unary/cast nodes are deduplicated, and\nunreachable nodes are not included in compiled plans. Capture is limited to 512\nnodes.\n\n**Outputs and diagnostics.** Callbacks can return tensors or nested arrays of\ntensors, preserving array keys. Up to four adjacent independent outputs with the\nsame shape can share a kernel. Other outputs use separate kernels; fusion does not\npromise one kernel for an entire callback. `getPlan()` reports step kinds, output\ncounts and the reason for each materialization, including zero-copy `view` steps\nfor layouts consumed by compatible native operations. `getStats()` exposes planned\nfused kernels, boundary steps, intermediate buffer count, scratch reuse and\nsuccessful replay count.\n\n**Capture limits.** During `run()`, CPU reads, legacy slicing (`__invoke()` and\n`[]`), serialization and operations not captured by the planner materialize their\nrequired inputs and continue eager execution; later elementwise operations can form\na new segment. During `compile()`, reads of placeholder data are rejected rather\nthan specializing on example values. Shape, stride and dtype queries do not execute\nkernels. Nested capture, Fiber switching, tensor mutation (including compound\nassignments), custom kernel launches and device changes/reset are prohibited during\ncapture. Exceptions restore eager execution; escaped tensors from an aborted capture\ncannot be read. `run()` remains synchronous.\n\n**PTX cache.** Generated PTX is cached per PHP request/thread, with LRU eviction at\n16 entries or 16 MiB. Keys include generated source (operations, constants, dtypes\nand layouts), compute capability and CUDA driver/runtime versions, not input\npointers. `Fusion::getCacheStats()` reports hits, misses, compilations and\nevictions. `Fusion::clearCache()` drops PTX without invalidating existing graphs.\n`compile()` also avoids repeated planning and module loading during replay.\n\nPlans containing generated kernels, matmul and reductions run on a private nonblocking stream. Reduction descriptors are kernel parameters rather than shared global state, so concurrent replays cannot overwrite each other's shapes.\n\n``` php\n$graph = Fusion::compile(\n    fn($a, $b, $c) => ($a + $b * $c)->sqrt(),\n    [$tensorA, $tensorB, $tensorC],\n    cudaGraph: true\n);\n$pending = $graph->runAsync($tensorA, $tensorB, $tensorC);\n$finished = $pending->isFinished(); // Query without waiting.\n$result = $pending->wait();        // Synchronize and retrieve outputs.\n```\n\n`cudaGraph: true` opts into a CUDA Graph executable for compatible plans. Kernel\nparameters are updated for new input/output pointers before each launch.\n`getStats()['backend']` is `cuda-graph`, `stream` or `native`.\n\n| Plan contains | Stream and `runAsync()` | CUDA Graph | \n|---|---|---|\n| Generated kernels only | ✅ | ✅ | \n| Matmul / reductions | ✅ | not yet | \n| Power boundaries | ❌ (synchronous native executor) | ❌ | \n\nCheck `cudaGraphCompatible` and `cudaGraphIncompatibility` to distinguish graph\nsupport from `asyncCompatible`. Power boundaries report the reason in\n`incompatibility` and reject `runAsync()`. CUDA failures are reported rather than\nhidden behind fallback.\n\n## Concurrency rules, scratch memory and profiling\n\nSynchronous compiled replay keeps a reusable stream, readiness event and scratch workspace; async replays use separate scratch storage. Slots are planned by last consumer, including native consumers and aliased views. Returned outputs always have independent storage across replays. Scratch remains allocated until the graph is released and counts toward the configured GPU memory budget.\n\nStream plans allow concurrent `runAsync()` calls. A CUDA Graph executable allows\nonly one outstanding replay: call `wait()` (or release the execution) before\nreplaying that graph again. A completion query alone does not release the\nexecution's retained resources. Inputs and closure-captured tensors remain alive\nuntil completion is collected. While an execution is outstanding, tensor mutation,\ncustom kernel launches and device changes/reset are blocked. Inputs must be ready\nbefore submission; independent custom async producer streams still require their\nexisting synchronization contract.\n\nRepeated `wait()` calls return the same outputs. Destroying a pending\n`FusionExecution` synchronizes before releasing storage. Execution submission\npreallocates tensors and can incur allocation synchronization; asynchronous kernel\nsubmission does not imply a zero-blocking PHP call. Device shape/stride metadata is\nallocated lazily, only for kernels that need it. Generated kernels, reductions,\nunaries and matmul use compiled or host descriptors instead of uploading metadata\nfor every result tensor.\n\nOptional synchronous replay phase timing:\n\n``` php\n$graph->setProfiling(true); // Enable timing and reset phase totals.\n$result = $graph->run($tensorA, $tensorB, $tensorC);\nprint_r($graph->getStats());\n$graph->setProfiling(false); // Disable timing and reset phase totals.\n```\n\n`profiledExecutions`, `bindTimeNs`, `executeTimeNs` and `collectTimeNs` accumulate\nsuccessful synchronous `run()` calls only. Execution timing includes preparation,\nsubmission and waiting; it is not isolated GPU kernel time. Profiling is off by\ndefault. `tensorAllocations` counts replay-created data tensors, not pool cache\nmisses or CUDA allocation calls. `workspaceBuffers` counts retained scratch slots;\n`bufferReuses` and `synchronizations` are cumulative replay counters.\n\nRun the focused comparison of eager, cached scoped, stream replay and CUDA Graph replay with:\n\n```\nphp -n -d extension=./cuda_build-8.3/modules/cuda.so \\\n  examples/07_fusion_graph.php --benchmark --elements=65536 --iterations=100\n```\n\nThe example validates output bytes, reports cold/cache compilation costs and end-to-end timings, and does not assume CUDA Graph is faster for every workload.\n\n[`examples/08_gpu_classifier.php`](https://github.com/lcmialichi/php-gpu-tensors/blob/main/examples/08_gpu_classifier.php)trains a deep MLP classifier on real datasets (MNIST, Fashion-MNIST, or CSV) without custom CUDA source or environment switches:\n\n- Parses and normalizes dataset files directly in PHP.\n- Packed float32 batches are uploaded once with `fromBuffer()` and kept on the GPU, including their transpose views.\n- Forward pass, numerically stable softmax cross-entropy, backward pass, and optimizer (AdamW or SGD) steps are compiled once per batch shape and replayed with new parameter tensors\n- Matmul/reduction boundaries and generated elementwise kernels share a private stream, with minimal synchronization.\n- Loss is transferred only on reporting epochs, outputting a colorful CLI report including macro-F1 score and confusion analysis\n\n```\nphp -n -d extension=./cuda_build-8.3/modules/cuda.so examples/08_gpu_classifier.php \\\n  --dataset=mnist --epochs=40 --batch-size=128 --optimizer=adam\n```\n\nDefaults are MNIST, 512-256 hidden layers, 40 epochs, batch size 128, and the AdamW optimizer. Every default run trains from deterministic initial weights. Omit `--no-save` to save parameters after successful evaluation; `--load` explicitly evaluates a compatible saved model instead of training. Dataset and model files are saved to the script's origin directory and are ignored by Git. The script reports one-time uploads, compilation, training throughput, and detailed evaluation metrics. Add `--profile` to print per-plan timing, allocation, scratch, and synchronization counters after training, or `--predict-index=N` to showcase inference on a specific sample.\n\n| API | Purpose | \n|---|---|\n| `Cuda\\CudaArray` | GPU allocation, tensor math, reductions, views, imports and `where()` | \n| `Cuda\\HostArray` /`Cuda\\ContiguousArray` | CPU storage, packed buffers, optional pinned memory and `toGpu()` | \n| `Cuda\\Fusion` /`Cuda\\FusionGraph` | Optional expression capture, compiled replay, PTX cache and plan diagnostics | \n| `Cuda\\FusionExecution` | Pending compatible execution, completion query and synchronized result collection | \n| `Cuda\\Compiler` /`Cuda\\CompiledModule` | NVRTC compilation, cached PTX, synchronous and asynchronous kernels | \n| `cuda_get_device_count()` and other`cuda_*` functions | Device selection, properties, memory and synchronization | \n| `Cuda\\Exception` | Base class for runtime, argument, allocation and compilation errors | \n\nThe annotated signatures are in [class stubs](https://github.com/lcmialichi/php-gpu-tensors/blob/main/stubs/cuda.stub.php) and\n[device function stubs](https://github.com/lcmialichi/php-gpu-tensors/blob/main/stubs/cuda_methods.stub.php); runnable examples live in\n[examples](https://github.com/lcmialichi/php-gpu-tensors/blob/main/examples/README.md).\n\n**Current limits:**\n\n- NVIDIA GPUs only, on Linux. GPU data has no CPU fallback.\n- No automatic differentiation; backpropagation is written explicitly.\n- `astype()` supports safe dtype conversions.\n- The core API baseline is frozen at `0.1.0` , and the current release is beta.\n- Matmul/reduction plans support async replay, but not CUDA Graph yet; power boundaries still require synchronous execution.\n\nThe core API has a frozen `0.1.0` baseline; eager execution remains the default\nand Fusion is explicitly opt-in. Build and CPU checks do not substitute for GPU\nruntime validation on each PHP version and thread mode.\n\n| PHP / mode | GPU | What was validated | \n|---|---|---|\n| 8.3 NTS | GeForce MX570 A | Current release: 41 GPU tests (including gradient/lifetime regressions) plus an Optdigits training check | \n| 8.5 NTS and ZTS, 8.1 ZTS | RTX A2000 | Earlier tensor/JIT validation only; does not cover the latest Fusion changes | \n| All ten PHP 8.1–8.5 × NTS/ZTS combinations | none | Builds and CPU/API checks for this release | \n\n**Tested on another GPU, PHP version or CUDA version?** Reports are very welcome:\nplease [open an issue](https://github.com/lcmialichi/php-gpu-tensors/issues) with\nyour PHP version, thread mode, GPU, driver and CUDA versions, and whether\n`./run-tests.sh --require-gpu` passed.\n\nContributions can be code, tests, documentation, runnable examples, or reports from\nanother PHP/CUDA/GPU combination. Browse\n[open issues](https://github.com/lcmialichi/php-gpu-tensors/issues),\n[report a bug](https://github.com/lcmialichi/php-gpu-tensors/issues/new?template=bug_report.yml),\nor [propose a feature](https://github.com/lcmialichi/php-gpu-tensors/issues/new?template=feature_request.yml).\nStart with the [contribution guide](https://github.com/lcmialichi/php-gpu-tensors/blob/main/CONTRIBUTING.md); documentation, examples, and\nhost-side tests are possible without an NVIDIA GPU.\n\nFor possible directions, see the [roadmap](https://github.com/lcmialichi/php-gpu-tensors/blob/main/ROADMAP.md). A feature does not need to\nbe listed there to be worth discussing.\n\nThe benchmark suite is maintained in the separate\n[PHP GPU Tensors Benchmarks repository](https://github.com/lcmialichi/php-gpu-tensors-benchmarks).\nIt includes focused `--matmul`, `--import` and `--fusion` runs, JSON/HTML reports,\nand a [beta.4 full-suite report](https://github.com/lcmialichi/php-gpu-tensors-benchmarks/blob/main/published-reports/beta4-php83-mx570/README.md)\ncovering 398 cases on PHP 8.3 NTS and an MX570 A, including Fusion replay and\nslicing. Earlier results include a\n[published PHP 8.5 NTS vs ZTS comparison](https://github.com/lcmialichi/php-gpu-tensors-benchmarks/blob/main/published-reports/php85-nts-vs-zts/README.md),\nas well as a [full-suite benchmark report](https://github.com/lcmialichi/php-gpu-tensors-benchmarks/blob/main/published-reports/php85-nts-vs-zts-full/README.md)\ncovering 368 cases across five workload groups with downloadable raw reports.\n\nLicensed under the [MIT License](https://github.com/lcmialichi/php-gpu-tensors/blob/main/LICENSE).", "url": "https://wpnews.pro/news/show-hn-php-gpu-tensors-native-gpu-operations-in-php", "canonical_source": "https://github.com/lcmialichi/php-gpu-tensors", "published_at": "2026-10-05 12:32:55+00:00", "updated_at": "2026-10-05 12:49:30.584177+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-infrastructure", "developer-tools"], "entities": ["PHP-GPU-tensors", "PHP", "NVIDIA CUDA", "NVRTC", "NVIDIA GeForce MX570 A", "PatchCamelyon", "lcmialichi", "CudaArray"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-php-gpu-tensors-native-gpu-operations-in-php", "markdown": "https://wpnews.pro/news/show-hn-php-gpu-tensors-native-gpu-operations-in-php.md", "text": "https://wpnews.pro/news/show-hn-php-gpu-tensors-native-gpu-operations-in-php.txt", "jsonld": "https://wpnews.pro/news/show-hn-php-gpu-tensors-native-gpu-operations-in-php.jsonld"}}