GPU tensors, runtime-compiled CUDA kernels and fused tensor expressions for PHP. No Python runtime required.
Native PHP extension for GPU tensors and NVIDIA CUDA-accelerated numerical workloads. Build tensor operations and machine-learning data pipelines in PHP, move data explicitly between host and GPU, and compile custom CUDA C++ kernels at runtime with NVRTC. Optionally fuse tensor expressions into compiled plans that replay quickly, with asynchronous execution and CUDA Graph for compatible plans.
Status: beta (0.1.0-beta.4) · Linux · PHP 8.1–8.5, NTS and ZTS · NVIDIA GPU
required. See Validation status for exactly what has been
tested on real GPUs.
- GPU tensors.
CudaArraysupports arithmetic, broadcasting, comparisons,matmul()(including batches), reductions, views and Python-styleslice(). - Explicit data movement. Packed buffers,
.npyimport and optional pinned host memory keep transfers cheap and visible. - Custom CUDA kernels. Compile CUDA C++ at runtime with NVRTC and launch it from PHP, synchronously or asynchronously.
- Optional kernel fusion. Capture a PHP closure once, then replay it as fused GPU kernels. In a small training workload this was about 6× faster than eager execution with identical results (seeResults ).
- Clear scope. A low-level GPU computing library, not a machine-learning framework: no automatic differentiation, no CPU fallback.
Contents: Quick look · Results · Install · Tensors · Slicing · Data pipelines · Custom kernels · Fusion · Streams and CUDA Graph · Training example · API and limits · Validation status · Contribute
use Cuda\CudaArray;
use Cuda\Fusion;
$a = CudaArray::ones([4]);
$b = CudaArray::full([4], 2.0);
$c = CudaArray::full([4], 3.0);
// Eager execution (default): one GPU operation per call.
$eager = $a + $b * $c;
// Fusion (opt-in): capture once, replay as fused GPU kernels.
$plan = Fusion::compile(fn($a, $b, $c) => $a + $b * $c, inputs: [$a, $b, $c]);
$fused = $plan->run($a, $b, $c);
print_r($fused->toArray()); // [7, 7, 7, 7]
Try a full training run on the GPU, written entirely in PHP:
php -n -d extension=./cuda_build-8.3/modules/cuda.so examples/08_gpu_classifier.php \
--epochs=200 --batch-size=256 --learning-rate=0.05 --no-save
Measured on PHP 8.3 NTS with an NVIDIA GeForce MX570 A (4 GB, compute capability 8.6), driver 12.6 and CUDA runtime 12.3. The models are small, so these numbers mostly show how much per-step overhead Fusion removes; they are not peak GPU throughput and not a general-purpose GPU benchmark.
Fusion vs. eager execution on the same model and data, with identical metrics (accuracy 77.29%, ROC AUC 0.853, same confusion matrix in both modes):
| Mode | Time per step | Patches per second | Training time (1,280 steps) |
|---|---|---|---|
| Eager | 4.08 ms | 125,444 | 5.22 s |
| Fusion (compiled replay) | 0.66 ms | 776,251 | 0.84 s |
Workload: a small MLP (hidden size 256) on 32,768 training and 4,096 test patches of PatchCamelyon (CC0), using 480 handcrafted features per patch (RGB mean/std over an 8×8 grid and its central 4×4), batch size 512, 20 epochs, learning rate 0.02. This is a performance demonstration, not a clinical model. The planner turned the 68 captured nodes of the training step into 11 fused kernels plus 9 native boundaries (matmul and reductions).
End-to-end training example (examples/08_gpu_classifier.php, a complete Multi-Layer Perceptron (MLP) for MNIST, Fashion-MNIST, or custom CSVs. It processes hundreds of thousands of samples per second using fused CUDA kernels, showcasing tensor operations, JIT compilation, math optimization (AdamW/SGD), and pure PHP data orchestration.
Full benchmark reports live in the benchmarks repository.
| Need | Details |
|---|---|
| GPU and driver | CUDA-capable NVIDIA GPU; the host driver provides libcuda.so.1 at runtime |
| Build toolchain | CUDA Toolkit (including NVRTC), C/C++ toolchain, make ,autoconf |
| PHP | 8.1–8.5 development headers ( phpize ,php-config ), NTS or ZTS |
| OS | Linux |
Building requires the toolkit; running requires the host driver.
git clone https://github.com/lcmialichi/php-gpu-tensors.git
cd php-gpu-tensors
./compile.sh
./run-tests.sh --require-gpu
php -n -d extension=./cuda_build-8.3/modules/cuda.so examples/01_basics_cuda_array.php
The build stays in cuda_build-<PHP major.minor>/modules/cuda.so; replace 8.3
above with the PHP version used by php-config. To install the extension and its
INI configuration instead, run ./compile.sh --install with permission to write
to your PHP extension/INI directories. To choose another PHP ABI:
PHP_BIN=php8.3 PHPIZE=phpize8.3 PHP_CONFIG=php-config8.3 ./compile.sh
PHP_BIN=php8.3 PHP_CONFIG=php-config8.3 ./run-tests.sh --require-gpu
Build options:
CUDA_HOMEif the toolkit is not at/usr/local/cuda.CUDA_ARCH=sm_86when cross-building.CUDA_USE_CUBLAS=no ./compile.shto use the extension's built-in CUDA matrix kernels instead of cuBLAS (cuBLAS is used when available for compatible larger matrix products).
The extension is published as
lcmialichi/php-gpu-tensors.
pie install lcmialichi/php-gpu-tensors:0.1.0-beta.4
Use 0.1.0-beta.4 or newer for the Fusion APIs and the training example
described here; older releases may not include them. PIE builds the native
extension for the selected PHP installation; it does not install an NVIDIA driver
or CUDA Toolkit. Those must already be available on the system, and GPU execution
additionally requires a compatible NVIDIA driver and a visible GPU.
Users with the NVIDIA Container Toolkit and a working host driver can build and test in the development image:
docker compose run --rm php_cuda_dev bash -lc './compile.sh && ./run-tests.sh --require-gpu'
./run-tests.sh runs CPU-side C tests and the PHP test suite. Use --cpu-only to
run only the host-side C tests in build environments without a GPU; this does not
validate CUDA execution. --require-gpu fails immediately when no GPU is visible,
instead of treating skipped GPU tests as success.
use Cuda\CudaArray;
$input = new CudaArray([[1, 2], [3, 4]], 'float32');
$weights = CudaArray::ones([2, 2]);
$output = $input->add($weights)->multiply(2);
$column_means = $input->mean(0);
print_r($output->toArray()); // [[4, 6], [8, 10]]
echo $output->dtype(); // float32
CudaArray holds GPU storage; operations return GPU tensors. toArray()
transfers to the CPU and expands all values into PHP arrays. Operations include
arithmetic, broadcasting, comparisons, matmul() (including batches), shape
views, and sum(), mean(), min(), max(), prod(), argMax(), and
argMin() with an optional axis. PHP arithmetic operators also dispatch to tensor
methods. mean() reduces all values when called without an axis, or reduces one
dimension when given an axis. It returns float32 for float32 input and
float64 for float64, integer, and boolean input.
CudaArray::slice() accepts a comma-separated expression, or one selector per
axis. Ranges have an exclusive stop, omitted bounds default to the axis limits,
and negative indices/bounds count from the end. Range bounds are clipped to the
axis size; individual indices outside the axis throw InvalidArgumentException.
$batch = $tensor->slice('10:20, :');
$columns = $tensor->slice(':, 2:8:2');
$lastRows = $tensor->slice('-10:');
$row = $tensor->slice(3); // Removes the first axis.
$oneRow = $tensor->slice('3:4'); // Keeps that axis with size 1.
// Dynamic equivalents: [start, stop] or [start, stop, step].
$batch = $tensor->slice([$start, $stop], null);
$columns = $tensor->slice(null, [2, 8, 2]);
Integers remove axes; null, ':', omitted axes and slice() with no arguments
select the full extent. Steps must be positive integers no larger than INT_MAX.
Zero/negative steps, ellipsis, new axes, index lists and executable expressions
are not supported. Expressions are parsed as integers/ranges, never evaluated as
PHP.
Views, empty ranges and legacy syntax #
A fully indexed tensor has shape [] and size 1; toArray() and
toHost()->toArray() represent it as a one-element array.
Outside Fusion, slices are zero-copy views with shared storage: writes through a view affect its parent, and the parent remains alive while the view is retained. Host transfers pack strided views into contiguous host storage. Row assignments involving strided tensors stage the source on the host before writing, preserving overlapping-source correctness; this is not a GPU-only bulk scatter operation. Fusion captures slice index transformations without materializing during capture; returned compiled/scoped outputs use the existing independent-output storage contract.
Empty ranges are supported, preserving the remaining shape: [0, 4] converts to
[], whereas [3, 0] converts to [[], [], []]. Elementwise operations, casts,
transfers, reshape/transpose and matmul handle zero elements. Sum/product over an
empty axis return 0/1; mean returns NaN. Min/max and arg reductions throw when an
empty reduced axis would produce values, because no identity/index exists. An
empty non-reduced output remains empty. Packed-buffer imports accept zero
dimensions; materialized empty results support serialization.
The legacy __invoke() and [] selection syntax remain unchanged, including
their inclusive ranges. Do not interpret their bounds as the new exclusive
slice() bounds.
Use this extension to build GPU-accelerated numerical steps into PHP
applications: tensor arithmetic, matrix multiplication, broadcasting, reductions
such as mean(), and custom CUDA kernels. These primitives support
machine-learning data preparation, inference and small training loops while the
data remains in NVIDIA GPU memory.
This is a low-level GPU computing library, not a complete machine-learning framework. The training example implements backpropagation explicitly; automatic differentiation and Python interoperability are not provided.
For data already in packed row-major bytes, avoid creating individual PHP
scalars. fromFile() reads raw bytes, whereas fromNpy() parses NumPy's .npy
format:
use Cuda\CudaArray;
use Cuda\HostArray;
$bytes = pack('g*', 1, 2, 3, 4); // little-endian float32
$host = HostArray::fromBuffer($bytes, [2, 2]);
$gpu = $host->toGpu();
$result = CudaArray::where(
CudaArray::fromBuffer(pack('C*', 1, 0), [2], 'bool'),
$gpu,
CudaArray::zeros([2, 2])
);
$cpuCopy = $result->toHost();
$raw = $cpuCopy->toBuffer();
HostArray is an alias of Cuda\ContiguousArray, a contiguous CPU tensor. Pass
pinned: true to its constructor or fromBuffer() for page-locked host storage
when repeated transfers justify the extra host memory. where() broadcasts its
three inputs; its mask treats nonzero values as true, and its two value tensors
must have the same dtype. .npy imports support C-order little-endian numeric and
boolean arrays; Fortran order and big-endian data are rejected.
Register the CUDA kernel's argument metadata, compile to PTX, then launch with explicit grid and block dimensions:
use Cuda\Compiler;
use Cuda\CudaArray;
$source = <<<'CUDA'
extern "C" __global__ void scale(float *data, int factor, int count)
{
int index = blockIdx.x * blockDim.x + threadIdx.x;
if (index < count) data[index] *= factor;
}
CUDA;
$compiler = new Compiler(source: $source);
$compiler->kernel('scale', [
['name' => 'data', 'type' => 'array', 'dtype' => 'float32'],
['name' => 'factor', 'dtype' => 'int32'],
['name' => 'count', 'dtype' => 'int32'],
]);
$module = $compiler->compile();
$module->initialize();
$data = CudaArray::ones([512]);
$module->launch('scale',
args: [$data, 3, 512],
config: ['block' => [256, 1, 1], 'grid' => [2, 1, 1]]
);
print_r(array_slice($data->toArray(), 0, 4)); // [3, 3, 3, 3]
launch() synchronizes; launchAsync() returns an operation ID for sync() or
wait(). Keep tensors alive until asynchronous work finishes. See
the JIT examples and
asynchronous execution.
Eager execution remains the default. Fusion is explicitly opt-in: existing code keeps using eager execution unless capture is enabled.
flowchart LR
A[PHP closure] --> B[Capture once<br/>metadata-only placeholders]
B --> C[Plan]
C --> D[Fused elementwise kernels<br/>NVRTC to PTX]
C --> E[Native boundaries<br/>matmul and reductions]
D --> F[Replay on a private stream]
E --> F
Fusion::run() captures tensor operations inside a callback and returns
materialized CudaArray outputs:
use Cuda\Fusion;
$result = Fusion::run(fn() => $tensorA + $tensorB * $tensorC);
$eager = Fusion::run(fn() => $tensorA + $tensorB * $tensorC, enabled: false);
For repeated execution, compile a specialized plan once:
$graph = Fusion::compile(
fn($a, $b, $c) => $a + $b * $c,
inputs: [$tensorA, $tensorB, $tensorC]
);
$result = $graph->run($tensorA, $tensorB, $tensorC);
$other = $graph->run($otherA, $otherB, $otherC);
print_r($graph->getStats());
print_r($graph->getPlan());
echo $graph->getSource();
What to know at a glance:
- Fused: addition, subtraction, multiplication, division, unary operations,
comparisons,
where(), safeastype()conversions and reshape/transpose/slice index transformations. - Execution boundaries: reductions,
matmul()(currently float32) and powers run on existing kernels; the fused kernels run in dependency order around them. - Replay: the PHP callback is not invoked again after compilation, and input values and pointers can change between executions.
- Inspection:
getStats(),getPlan()andgetSource()show what the planner produced; for$a + $b * $cthe plan has one fused kernel and no intermediate data buffers.
Fusion reference: compile and replay, fusion rules, capture limits and caching #
Compile and replay. Compilation invokes the callback once with metadata-only placeholders. Replay does not invoke PHP callback code again. Inputs must match the example shapes, dtypes and strides, and execution must use the compilation device and CUDA context. Input values and pointers can change between executions; keep the device/context alive until the graph is released. Tensors captured by a closure are retained by the graph; use callback parameters for replaceable inputs.
What fuses. Addition, subtraction, multiplication, division, unary operations
and comparison methods fuse into elementwise kernels, with broadcasting,
strided/view inputs, scalar operands and dtype promotion. Each node converts to its
own result dtype, preserving intermediate rounding/narrowing. Generated kernels use
the eager backend's fast-math settings but disable cross-node FMA contraction.
where(), safe explicit astype() conversions and reshape/transpose/slice index
transformations also fuse. Safe casts work in eager execution too, using the same
generated conversion kernel; unsafe narrowing retains the existing rejection
policy.
Boundaries and planning. Reductions, matmul() (currently float32) and powers
are execution boundaries using existing kernels. All generated elementwise kernels
in a plan are compiled together through the existing Compiler NVRTC
infrastructure, then executed in dependency order around these boundaries.
Expressions split at a weighted cost budget of 32 (math functions, index transforms
and float64 have higher cost). Shared expensive expressions can be materialized
instead of recomputed. Pure repeated binary/unary/cast nodes are deduplicated, and
unreachable nodes are not included in compiled plans. Capture is limited to 512
nodes.
Outputs and diagnostics. Callbacks can return tensors or nested arrays of
tensors, preserving array keys. Up to four adjacent independent outputs with the
same shape can share a kernel. Other outputs use separate kernels; fusion does not
promise one kernel for an entire callback. getPlan() reports step kinds, output
counts and the reason for each materialization, including zero-copy view steps
for layouts consumed by compatible native operations. getStats() exposes planned
fused kernels, boundary steps, intermediate buffer count, scratch reuse and
successful replay count.
Capture limits. During run(), CPU reads, legacy slicing (__invoke() and
[]), serialization and operations not captured by the planner materialize their
required inputs and continue eager execution; later elementwise operations can form
a new segment. During compile(), reads of placeholder data are rejected rather
than specializing on example values. Shape, stride and dtype queries do not execute
kernels. Nested capture, Fiber switching, tensor mutation (including compound
assignments), custom kernel launches and device changes/reset are prohibited during
capture. Exceptions restore eager execution; escaped tensors from an aborted capture
cannot be read. run() remains synchronous.
PTX cache. Generated PTX is cached per PHP request/thread, with LRU eviction at
16 entries or 16 MiB. Keys include generated source (operations, constants, dtypes
and layouts), compute capability and CUDA driver/runtime versions, not input
pointers. Fusion::getCacheStats() reports hits, misses, compilations and
evictions. Fusion::clearCache() drops PTX without invalidating existing graphs.
compile() also avoids repeated planning and module during replay.
Plans containing generated kernels, matmul and reductions run on a private nonblocking stream. Reduction descriptors are kernel parameters rather than shared global state, so concurrent replays cannot overwrite each other's shapes.
$graph = Fusion::compile(
fn($a, $b, $c) => ($a + $b * $c)->sqrt(),
[$tensorA, $tensorB, $tensorC],
cudaGraph: true
);
$pending = $graph->runAsync($tensorA, $tensorB, $tensorC);
$finished = $pending->isFinished(); // Query without waiting.
$result = $pending->wait(); // Synchronize and retrieve outputs.
cudaGraph: true opts into a CUDA Graph executable for compatible plans. Kernel
parameters are updated for new input/output pointers before each launch.
getStats()['backend'] is cuda-graph, stream or native.
| Plan contains | Stream and runAsync() |
CUDA Graph |
|---|---|---|
| Generated kernels only | ✅ | ✅ |
| Matmul / reductions | ✅ | not yet |
| Power boundaries | ❌ (synchronous native executor) | ❌ |
Check cudaGraphCompatible and cudaGraphIncompatibility to distinguish graph
support from asyncCompatible. Power boundaries report the reason in
incompatibility and reject runAsync(). CUDA failures are reported rather than
hidden behind fallback.
Concurrency rules, scratch memory and profiling #
Synchronous compiled replay keeps a reusable stream, readiness event and scratch workspace; async replays use separate scratch storage. Slots are planned by last consumer, including native consumers and aliased views. Returned outputs always have independent storage across replays. Scratch remains allocated until the graph is released and counts toward the configured GPU memory budget.
Stream plans allow concurrent runAsync() calls. A CUDA Graph executable allows
only one outstanding replay: call wait() (or release the execution) before
replaying that graph again. A completion query alone does not release the
execution's retained resources. Inputs and closure-captured tensors remain alive
until completion is collected. While an execution is outstanding, tensor mutation,
custom kernel launches and device changes/reset are blocked. Inputs must be ready
before submission; independent custom async producer streams still require their
existing synchronization contract.
Repeated wait() calls return the same outputs. Destroying a pending
FusionExecution synchronizes before releasing storage. Execution submission
preallocates tensors and can incur allocation synchronization; asynchronous kernel
submission does not imply a zero-blocking PHP call. Device shape/stride metadata is
allocated lazily, only for kernels that need it. Generated kernels, reductions,
unaries and matmul use compiled or host descriptors instead of up metadata
for every result tensor.
Optional synchronous replay phase timing:
$graph->setProfiling(true); // Enable timing and reset phase totals.
$result = $graph->run($tensorA, $tensorB, $tensorC);
print_r($graph->getStats());
$graph->setProfiling(false); // Disable timing and reset phase totals.
profiledExecutions, bindTimeNs, executeTimeNs and collectTimeNs accumulate
successful synchronous run() calls only. Execution timing includes preparation,
submission and waiting; it is not isolated GPU kernel time. Profiling is off by
default. tensorAllocations counts replay-created data tensors, not pool cache
misses or CUDA allocation calls. workspaceBuffers counts retained scratch slots;
bufferReuses and synchronizations are cumulative replay counters.
Run the focused comparison of eager, cached scoped, stream replay and CUDA Graph replay with:
php -n -d extension=./cuda_build-8.3/modules/cuda.so \
examples/07_fusion_graph.php --benchmark --elements=65536 --iterations=100
The example validates output bytes, reports cold/cache compilation costs and end-to-end timings, and does not assume CUDA Graph is faster for every workload.
examples/08_gpu_classifier.phptrains a deep MLP classifier on real datasets (MNIST, Fashion-MNIST, or CSV) without custom CUDA source or environment switches:
- Parses and normalizes dataset files directly in PHP.
- Packed float32 batches are uploaded once with
fromBuffer()and kept on the GPU, including their transpose views. - Forward pass, numerically stable softmax cross-entropy, backward pass, and optimizer (AdamW or SGD) steps are compiled once per batch shape and replayed with new parameter tensors
- Matmul/reduction boundaries and generated elementwise kernels share a private stream, with minimal synchronization.
- Loss is transferred only on reporting epochs, outputting a colorful CLI report including macro-F1 score and confusion analysis
php -n -d extension=./cuda_build-8.3/modules/cuda.so examples/08_gpu_classifier.php \
--dataset=mnist --epochs=40 --batch-size=128 --optimizer=adam
Defaults are MNIST, 512-256 hidden layers, 40 epochs, batch size 128, and the AdamW optimizer. Every default run trains from deterministic initial weights. Omit --no-save to save parameters after successful evaluation; --load explicitly evaluates a compatible saved model instead of training. Dataset and model files are saved to the script's origin directory and are ignored by Git. The script reports one-time uploads, compilation, training throughput, and detailed evaluation metrics. Add --profile to print per-plan timing, allocation, scratch, and synchronization counters after training, or --predict-index=N to showcase inference on a specific sample.
| API | Purpose |
|---|---|
Cuda\CudaArray |
GPU allocation, tensor math, reductions, views, imports and where() |
Cuda\HostArray /Cuda\ContiguousArray |
CPU storage, packed buffers, optional pinned memory and toGpu() |
Cuda\Fusion /Cuda\FusionGraph |
Optional expression capture, compiled replay, PTX cache and plan diagnostics |
Cuda\FusionExecution |
Pending compatible execution, completion query and synchronized result collection |
Cuda\Compiler /Cuda\CompiledModule |
NVRTC compilation, cached PTX, synchronous and asynchronous kernels |
cuda_get_device_count() and othercuda_* functions |
Device selection, properties, memory and synchronization |
Cuda\Exception |
Base class for runtime, argument, allocation and compilation errors |
The annotated signatures are in class stubs and device function stubs; runnable examples live in examples.
Current limits:
- NVIDIA GPUs only, on Linux. GPU data has no CPU fallback.
- No automatic differentiation; backpropagation is written explicitly.
astype()supports safe dtype conversions.- The core API baseline is frozen at
0.1.0, and the current release is beta. - Matmul/reduction plans support async replay, but not CUDA Graph yet; power boundaries still require synchronous execution.
The core API has a frozen 0.1.0 baseline; eager execution remains the default
and Fusion is explicitly opt-in. Build and CPU checks do not substitute for GPU
runtime validation on each PHP version and thread mode.
| PHP / mode | GPU | What was validated |
|---|---|---|
| 8.3 NTS | GeForce MX570 A | Current release: 41 GPU tests (including gradient/lifetime regressions) plus an Optdigits training check |
| 8.5 NTS and ZTS, 8.1 ZTS | RTX A2000 | Earlier tensor/JIT validation only; does not cover the latest Fusion changes |
| All ten PHP 8.1–8.5 × NTS/ZTS combinations | none | Builds and CPU/API checks for this release |
Tested on another GPU, PHP version or CUDA version? Reports are very welcome:
please open an issue with
your PHP version, thread mode, GPU, driver and CUDA versions, and whether
./run-tests.sh --require-gpu passed.
Contributions can be code, tests, documentation, runnable examples, or reports from another PHP/CUDA/GPU combination. Browse open issues, report a bug, or propose a feature. Start with the contribution guide; documentation, examples, and host-side tests are possible without an NVIDIA GPU.
For possible directions, see the roadmap. A feature does not need to be listed there to be worth discussing.
The benchmark suite is maintained in the separate
PHP GPU Tensors Benchmarks repository.
It includes focused --matmul, --import and --fusion runs, JSON/HTML reports,
and a beta.4 full-suite report
covering 398 cases on PHP 8.3 NTS and an MX570 A, including Fusion replay and
slicing. Earlier results include a
published PHP 8.5 NTS vs ZTS comparison,
as well as a full-suite benchmark report
covering 368 cases across five workload groups with downloadable raw reports.
Licensed under the MIT License.