Discover Arm contributions to LLVM 23, including architecture support, performance improvements, code generation, tooling, libraries, and MLIR updates
LLVM 23.1.0 was released on August 26, 2026. Teams across Arm contributed almost 1300 patches to improve architecture support, libraries, tools, performance and increasingly ML workloads. This post provides an overview of the key contributions.
To find out more about the previous LLVM release, you can read What is new in LLVM 22?
New architecture and CPU support #
Architecture support
By Maciej Gabka
LLVM 23 improves architecture support and Arm C Language Extensions (ACLE) support for Arm developers. It also improves architecture correctness in the AArch64 assembler and disassembler, aligning LLVM with the June 2026 updates to the Arm A-profile A64 Instruction Set Architecture**.** In addition, LLVM 23 adds assembly support for features such as the Extended Hint instruction space (FEAT_HINTE).
LLVM 23 provides support for the beta ACLE specification for data processing features added as part of Armv9.6-A architecture, including SVE2p2 and SME2p2, and alpha ACLE specification for Armv9.7-A data-processing features, including SVE2p3 and SME2p3. Clang now defines the appropriate ACLE feature-test macros when targeting Armv9.6-A or Armv9.7-A, allowing developers to select architecture-specific code paths at compile time.
LLVM 23 improves FP8 ACLE code generation by modeling the FPMR registers as a target-specific memory location. This enables more precise dependency and alias analysis, allowing redundant writes to be eliminated and loop-invariant writes to be safely hoisted out of loops.
CPU support
LLVM 23 added support for the Arm AGI CPU, which is the first production silicon developed by Arm, based on Armv9.2-A and designed specifically for AI agentic workloads. Read more about it in Introducing Arm AGI CPU.
LLVM 23 performance benefited from scheduling models for C1-Nano, C1-Premium and C1-Ultra CPUs.
Performance improvements #
Arm AGI CPU improvements
By Biplob Mishra
LLVM 23 improves tuning for the Arm AGI CPU by adjusting the interleave factor used for vector, scalar, and reduction loops to better match the characteristics of this CPU. These improvements provide measurable gains from native CPU tuning compared with generic tuning. On Arm AGI CPU, native tuning improves performance by 1.7% on average on SPEC CPU2017 and 0.8% on average on SPEC CPU2026.
SPEC Improvements
By Kiran Chandramohan
Arm teams contributed to several improvements in LLVM 23 that boost performance across SPEC benchmarks. In addition to developing new optimizations, the teams continued to monitor SPEC performance throughout the LLVM 23 development cycle and fixed several performance regressions, helping maintain or improve overall performance compared with LLVM 22.
Some notable highlights:
- 777.zstd_r – 4% improvement: Changes to
AggressiveInstCombineallow LLVM to recognize split-width 32-bitcttz/ctlzpatterns and replace them with wider 64-bit intrinsics, enabling further optimization. - 541.leela_r – 3% improvement: Changes to the AArch64 cost model for vector shifts with non-uniform constant shift amounts enable better vectorization decisions. This work originated from the investigation of a regression in
imagickand also resulted in a significant improvement forleela. - 753.ns3_r – 1% improvement: Changes to
SimplifyCFGexposed additional simplification opportunities by identifying more cases where PHI incoming values lead to undefined behavior.
Code generation improvements #
Extending the capabilities of partial reductions
By Sander De Smalen
Partial reductions let the vectorizer accumulate into a wider type without first widening every input element. For dot products and similar kernels, this reduces the number of explicit extensions, lowers register pressure, and enables the use of dedicated instructions.
LLVM 23 supports many more of these patterns, including floating-point reductions. A half-precision dot product written in ordinary C++ can now make use of SVE fdot when using -ffast-math:
float dot(const __fp16 *a, const __fp16 *b, int n) {
float sum = 0.0f;
for (int i = 0; i < n; ++i)
sum += (float)a[i] * (float)b[i];
return sum;
}
The resulting vector loop uses the more efficient fdot instruction, followed by a horizontal reduction:
.Lvector_body:
ld1h { z1.h }, p0/z, [x0, x10, lsl #1]
ld1h { z2.h }, p0/z, [x1, x10, lsl #1]
inch x10
cmp x9, x10
fdot z0.s, z2.h, z1.h
b.ne .Lvector_body
ptrue p0.s
faddv s0, p0, z0.s
See the full Compiler Explorer example.
Partial reductions now also work when the operation is predicated inside the loop. Code such as:
int64_t masked_dot(const int16_t *a, const int16_t *b,
const int16_t *mask, int n) {
int64_t sum = 0;
for (int i = 0; i < n; ++i)
if (mask[i] > 0)
sum += (int64_t)a[i] * b[i];
return sum;
}
can stay vectorized while ensuring inactive lanes do not contribute to the result. LLVM creates predicates from the mask, uses them for the input loads, and performs the partial reduction with sdot:
ld1h { z2.h }, p0/z, [x2, x10, lsl #1]
cmpgt p1.h, p0/z, z2.h, #0
ld1h { z2.h }, p1/z, [x8, x10, lsl #1]
ld1h { z3.h }, p1/z, [x1, x10, lsl #1]
inch x10
sel z2.h, p1, z2.h, z0.h
sdot z1.d, z3.h, z2.h
The before-and-after output shows the full generated function.
The matching is no longer limited to a single, canonical sum += product form. LLVM 23 can handle subtraction performed in the middle block:
for (int i = 0; i < n; ++i)
sum -= (int64_t)a[i] * b[i];
The vector loop accumulates the products with sdot; the subtraction from the initial value is then performed in the middle block:
.Lvector_body:
ld1h { z1.h }, p0/z, [x0, x10, lsl #1]
ld1h { z2.h }, p0/z, [x1, x10, lsl #1]
inch x10
sdot z0.d, z2.h, z1.h
cmp x9, x10
b.ne .Lvector_body
ptrue p0.d
uaddv d0, p0, z0.d
fmov x10, d0
sub x2, x2, x10
and mixed add/subtract chains:
for (int i = 0; i < n; ++i) {
sum += (int64_t)a[i] * b[i];
sum -= (int64_t)c[i] * d[i];
}
In the generated SVE loop, both products use sdot. The subr operations allow the additions and subtractions to be represented in the same partial-reduction chain:
ld1h { z1.h }, p0/z, [x0, x9, lsl #1]
ld1h { z2.h }, p0/z, [x1, x9, lsl #1]
sdot z0.d, z2.h, z1.h
ld1h { z1.h }, p0/z, [x2, x9, lsl #1]
ld1h { z2.h }, p0/z, [x3, x9, lsl #1]
subr z0.d, z0.d, #0
sdot z0.d, z2.h, z1.h
subr z0.d, z0.d, #0
Explore the generated code for the subtract reduction and the add/sub chain.
Support has also been added for absolute-difference reductions, a common building block in image, video, and signal-processing code:
unsigned absdiff(const unsigned char *a, const unsigned char *b, int n) {
unsigned sum = 0;
for (int i = 0; i < n; ++i)
sum += __builtin_abs((int)a[i] - (int)b[i]);
return sum;
}
LLVM can now recognize the widening absolute difference as part of a partial reduction instead of materializing a less efficient sequence. In this example, uabd calculates the byte-wise absolute differences and udot with a vector of ones accumulates them into 32-bit lanes:
.Lvector_body:
ldr z2, [x12]
ldr z3, [x13]
incb x13
incb x12
subs x14, x14, x8
uabd z2.b, p0/m, z2.b, z3.b
udot z0.s, z2.b, z1.b
b.ne .Lvector_body
ptrue p0.s
uaddv d0, p0, z0.s
fmov w8, s0
See the full Compiler Explorer example.
Vectorizing find-last reductions
By Sander De Smalen
Conditional scalar assignments inside a loop often encode a “find last” operation:
int last_match(const int *cond, const int *values, int n, int needle) {
int result = -1;
for (int i = 0; i < n; ++i)
if (cond[i] == needle)
result = values[i];
return result;
}
These loops are reductions, even though they do not look like a sum or a minimum. LLVM 23 has better support for recognizing and vectorizing these conditional assignments as find-last reductions. The vector loop compares several elements at once and conditionally retains their indices:
.Lvector_body:
ld1w { z2.s }, p1/z, [x0, x8, lsl #2] // load cond[i]
cmpeq p2.s, p1/z, z2.s, z1.s // cond[i] == needle
csetm x11, ne
whilelo p3.s, xzr, x11
ld1w { z2.s }, p2/z, [x1, x8, lsl #2] // load values[i]
incw x8
mov p0.b, p3/m, p2.b // if any cond[i] is true, then copy cond[] to p0
mov z0.s, p3/m, z2.s // if any cond[i] is true, then copy values[] to z0
cmp x10, x8
b.ne .Lvector_body
mov w8, #-1
cmp x10, x9
clastb w8, p0, w8, z0.s // extract last active lane, from values[]
This pattern also appears in parsers, scans, and loops that retain the most recent qualifying value.
See the full Compiler Explorer example.
Explicit floating-point reduction builtins
By Sander De Smalen
Clang now provides two floating-point reduction builtins:
float ordered = __builtin_reduce_in_order_fadd(v, start);
float fast = __builtin_reduce_assoc_fadd(v, start);
The optional second argument supplies the starting value; without it, the reduction starts from negative zero. The names make the key semantic choice explicit:
__builtin_reduce_in_order_faddevaluates the reduction in lane order.__builtin_reduce_assoc_faddallows reassociation, giving the optimizer freedom to use a tree reduction or target-specific instructions.
This allows programmers to specify the required reduction semantics without applying broad fast-math options to the rest of the translation unit.
Improved interleaving and SVE shuffles
By Sander De Smalen
Interleaved data is common in pixels, complex values, structures of channels, and packed sensor input. A lower vectorization factor may be appropriate for a loop, but previously it could prevent LLVM from selecting an interleaved memory operation.
LLVM 23 allows interleave shuffles at these lower vectorization factors for SVE, including cases where a tbl sequence efficiently combines deinterleaving with zero-extension. In suitable kernels, one table lookup per output replaces a longer sequence of unpack and shuffle instructions:
struct Pixel { unsigned short blue, green, red, alpha; };
void sum_channels(const Pixel *__restrict pixels,
double *__restrict sums, unsigned long count) {
double blue = 0, green = 0, red = 0, alpha = 0;
for (unsigned long i = 0; i < count; ++i) {
blue += pixels[i].blue;
green += pixels[i].green;
red += pixels[i].red;
alpha += pixels[i].alpha;
}
sums[0] = blue; sums[1] = green;
sums[2] = red; sums[3] = alpha;
}
Here, converting 16-bit channels to double makes the natural SVE vectorization factor smaller than the four-channel interleave factor. With -O3 -ffast-math -march=armv9-a+sve2, LLVM vectorizes the loop at vscale x 2 and folds both the deinterleave and zero-extension into the table lookups:
tbl z8.h, { z28.h }, z24.h
tbl z9.h, { z29.h }, z24.h
ucvtf z8.d, p0/m, z8.d
ucvtf z9.d, p0/m, z9.d
See the full Compiler Explorer example.
Improved multi-vector loads and stores
By Sander De Smalen
SVE and SME provide multi-vector memory operations. LLVM 23 improves the use of the instruction's addressing modes. For example,
#include <arm_sve.h>
void copy_vector_pair(unsigned char *dst, const unsigned char *src) {
svcount_t all = svptrue_c8();
svuint8x2_t data = svld1_vnum_u8_x2(all, src, -2);
svst1_vnum_u8_x2(all, dst, 14, data);
}
LLVM 23 folds both offsets directly into the multi-vector instructions, avoiding separate pointer adjustments:
ptrue pn8.b
ld1b { z0.b, z1.b }, pn8/z, [x1, #-2, mul vl]
st1b { z0.b, z1.b }, pn8, [x0, #14, mul vl]
See the full Compiler Explorer example
Streaming-mode for SME2 code also benefits from enabling sub-register liveness. For example, this kernel loads two rows as vector pairs, then transposes the tuple view for a dot product:
#include <arm_sme.h>
void dot_two_rows(unsigned long stride, const unsigned char *ptr)
__arm_streaming __arm_inout("za") {
svcount_t all = svptrue_c8();
svuint8x2_t row0 = svld1_u8_x2(all, ptr);
svuint8x2_t row1 = svld1_u8_x2(all, ptr + stride);
svuint8x2_t lhs = svcreate2_u8(svget2_u8(row0, 0), svget2_u8(row1, 0));
svuint8x2_t rhs = svcreate2_u8(svget2_u8(row0, 1), svget2_u8(row1, 1));
svdot_za32_u8_vg1x2(0, lhs, rhs);
}
LLVM 22 selects strided multi-vector loads, but must copy half of each tuple into consecutive registers for udot:
ld1b { z16.b, z24.b }, pn8/z, [x1]
ld1b { z17.b, z25.b }, pn8/z, [x1, x0]
mov z0.d, z24.d
mov z1.d, z25.d
udot za.s[w8, 0, vgx2], { z16.b, z17.b }, { z0.b, z1.b }
LLVM 23 can keep both tuple views live and removes the copies:
ld1b { z16.b, z24.b }, pn8/z, [x1]
ld1b { z17.b, z25.b }, pn8/z, [x1, x0]
udot za.s[w8, 0, vgx2], { z16.b, z17.b }, { z24.b, z25.b }
See the full Compiler Explorer example
Improved use of SVE immediates
By Sander De Smalen
LLVM 23 recognizes more opportunities to use SVE instructions that have an immediate form, instead of first materializing a constant in a vector register and using an equivalent NEON instruction that lacks an immediate form. That saves an instruction and, just as importantly, one temporary register.
For example:
#include <arm_neon.h>
uint32x4_t increment(uint32x4_t x) {
return vaddq_u32(x, vdupq_n_u32(1));
}
LLVM 22 materialized the splat in a second vector register before using a NEON add:
movi v1.4s, #1
add v0.4s, v0.4s, v1.4s
LLVM 23 can instead use the SVE immediate form directly:
add z0.s, z0.s, #1
See the Compiler Explorer example
Additional instruction-selection improvements
By Sander De Smalen
LLVM 23 also includes local AArch64 combines that remove redundant instructions.
One example folds a condition-code materialization followed by a bit-test branch:
cset w8, eq
tbnz w8, #0, .Lmatch
directly into the equivalent conditional branch:
b.eq .Lmatch
Although individually small, these folds can reduce code size and leave fewer instructions for later scheduling stages.
Improved carry-less multiplication
By Matthew Devereau
LLVM 23 improves AArch64 code generation for carry-less multiplication, an operation commonly used in cryptography and data-integrity algorithms. More fixed-width and scalable-vector forms are now mapped to sequences that utilize PMUL/ PMULL instructions when SVE2/SME and AES extensions are available and scalar forms can now utilize the vector register. The cost model has also been updated so that LLVM can make better optimization decisions.
Faster compilation with GlobalISel
By Cullen Rhodes
Compile-time has improved significantly in LLVM 23. Debug builds (-O0 -g) are 15% faster on CTMark, with some workloads such as SQLite almost 25% faster. Optimized builds (-O3) are also 8.7% faster. This is largely due to extensive improvements in GlobalISel, the default instruction selector at -O0, as well as broader improvements in Clang debug info and LLVM code generation.
Tools improvements #
BOLT improvements
By Paschalis Mpeis
LLVM 23 improves BOLT support for AArch64 binaries, with contributions focused on Branch Target Identification (BTI), compare-and-branch instructions, long-jump code layout, and usability.
We introduced initial BOLT support for processing BTI-enabled AArch64 binaries. BOLT can now check and patch PLT entries and indirect branch targets where possible. It can disassemble PLT entries in BTI binaries, patch LLD-generated PLTs with BTI landing pads, and patch ignored functions in place when indirect branches target them.
We also added compact-code-model support for Armv9.6-A FEAT_CMPBR compare-and-branch instructions. This includes support for block reordering, function splitting, branch inversion where legal, and trampoline handling for difficult branch cases.
This release also improves AArch64 usability. Unsupported or target-specific BOLT passes now report clear errors instead of crashing or silently doing nothing, and new documentation records AArch64 optimization flag support. BOLT also gains support for SHT_CREL code relocations, reports compressed debug sections cleanly, improves --hugify long-jump layout handling, and expands Arm test coverage.
Memory Protection Keys and Permission Overlays support in LLDB
By David Spickett
Memory permissions are typically set in the page tables. Programs often change these permissions at runtime. For example they might initialise key structures and then make them read-only to protect them against attackers trying to modify them later.
However, changing page table permissions is expensive. On Linux it is usually done with the mprotect system call. Linux has an alternative called "Memory Protection Keys" which allows permissions to be changed without a system call. These keys work together with a hardware feature which stores "permission overlays". On AArch64 this is the Permission Overlay Extension (FEAT_S1POE) introduced in Armv8.8.
New in LLDB 23 is support for debugging Memory Protection Keys and permission overlays on AArch64 Linux.
Below is an example of a protection key fault:
(lldb) c
Process 462 resuming
Process 462 stopped
<...> stop reason = signal SIGSEGV: failed protection key checks (fault address=0xffffff7d60000)
<...>
-> 106 read_only_page[0] = '?';
This program tried to write to memory and was not allowed to do so. You can tell that the page table entry did have write permissions because of the description of the SIGSEGV signal. It refers specifically to protection keys (SEGV_PKUERR) , as opposed to "invalid permissions for mapped object" (SEGV_ACCERR) which is generated when the page table permissions are the cause.
You can check the permissions with the memory region command:
(lldb) memory region read_only_page
[0x000ffffff7d60000-0x000ffffff7d70000) rw-
protection key: 6 (r--, effective: r--)
In prior versions of LLDB, all you would see is the rw- on the 2nd line. This represents the permissions in the page table. They allow reading (r) and writing ( w), but not execution ( x).
New in LLDB 23 is the 3rd line. It shows:
- The protection key assigned to this memory mapping (
6). - The permission overlay that the key refers to (
r--). - The effective permissions (
r--)after combining the page table permissions with the permission overlay.
Permission overlays can only remove permissions. In this example rw- was overlaid with r-- . This overlay keeps the r permission but removes the w and x permissions. That is why the program failed to write. This memory mapping is read-only.
LLDB also lets you access the POR_EL0 register where the permission overlays are stored. To read about this and other AArch64 Linux features, see the LLDB documentation.
Flang improvements
By Tom Eccles
Arm's contributions to LLVM 23 improved Flang's compatibility, performance and OpenMP support. Taskloop lowering is now more robust, with better handling of loop bounds, steps and privatized character values. Flang also supports the non-standard RTC intrinsic used by some existing Fortran applications. A new opt-in mode (-freal-sum-association in LLVM 23 or -ffp-sum-association in LLVM 24) can reassociate real-valued sums where permitted by the Fortran standard, exposing additional optimization opportunities while preserving explicit parentheses. Our biggest contribution was continuing upstream maintenance for flang's internal MLIR dialects, OpenMP on CPU support, and Windows support. I was the most active reviewer for flang OpenMP support (approximately 18% of reviews), and ranked third in flang as a whole (approximately 9% of reviews).
BOLT optimization for Flang
By Pawel Osmialowski
The purpose of BOLTing (optimizing with BOLT) the compiler itself is to make it compile the code faster; this is not about the optimization capabilities of the compiler itself (a common misconception).
Following the already existing methodology for having the Clang compiler BOLTed, we have contributed a similar method for BOLTing the Flang compiler. As with Clang, our approach is also based on the CMake caches and reuses the in-tree profiling utilities that were already used for BOLTing Clang. The proposed CMake caches introduce two optimization methods: BOLT and BOLT+PGO (with and without LTO). We did our performance evaluation of the BOLT+PGO method on different servers using three Fortran codebase examples (the DBCSR library, the Polyhedron benchmark, and the LAPACK library) and saw the following speedup in compilation times:
- DBCSR: 12.9 – 21.9%
- Polyhedron (pb11): 12.7 – 18.7%
- LAPACK-3.6.1: 15.3 – 16.2%
During this work, however, we also discovered that the PGO-based optimization process is affected by a race condition that can cause the performance-training stage to hang intermittently. Until this issue is resolved, we recommend using BOLT without PGO when optimizing Flang.
Windows on Arm and Arm64EC support
By David Truby
LLVM 23 extends Windows on Arm support across OpenMP, Flang and native interoperability. The OpenMP runtime now supports Arm64EC and Arm64X configurations, making it easier to combine Arm-native and emulated code. Flang’s runtime gains improved Windows implementations for date and time queries and file handling, together with Arm64EC build fixes. LLVM can also generate Arm64EC interoperability thunks for functions using bfloat16 values.
Libraries improvements #
Optimized AArch32 software floating point functions
By Simon Tatham
In the compiler-rt builtins library, there is now a full set of highly optimized implementations of floating-point basic arithmetic, for AArch32 systems supporting either of the Arm or Thumb2 instruction set but with no hardware FPU. A small number of functions also have optimized Thumb1 implementations, suitable for Armv6-M or Armv8-M Baseline. On average, the new functions run roughly twice as fast as the generic functions they replaced.
Libc vector math
By Dylan Fleming
LLVM 23 introduces a mathvec component into the LLVM libc, providing LLVM vector math implementations of standard math functions that can ship alongside their scalar counterparts. Current work has focused on establishing the component structure, testing infrastructure, and delivering proofs of concept for both a cross-architecture (generic) and an AdvSIMD-optimised single-precision exponential (expf).
Like scalar routines, vector routines are correctly rounded: they return the floating-point value nearest to the exact mathematical result for the supported rounding mode (presently, round to nearest, with ties to even). While this level of accuracy may be excessive for many use cases, correct rounding offers additional benefits.
Correctly rounded functions preserve properties of mathematical functions that lower-accuracy implementations may lose, potentially leading to unintuitive outputs, logical errors, and poor numerical stability. Among the properties frequently required by applications, monotonicity is preserved. For example, x <= y implies f(x) <= f(y) holds for floating-point quantities too. Perhaps more noticeably, functions remain within their true bounds, such as |sin(x)| <= 1 . Correct rounding also guarantees bitwise-identical results across:
- Scalar and vector implementations, enabling safe auto-vectorization without changing numerical behavior.
- Architectures and platforms, improving reproducibility and avoiding target-dependent production results.
- Library versions, helping users adopt the latest version and optimizations seamlessly.
Going forward, we plan to increase the coverage of vector routines, initially focusing on FP32 and FP16 before expanding to FP64 and BF16. These will first be delivered in generic form and enabled on all supported architectures, while target-specific optimizations will follow where observable gains can be achieved.
MLIR Improvements #
TOSA
By Luke Hutton
The TOSA MLIR dialect in LLVM 23 continues to support TOSA 1.0 while introducing early support for new features from the TOSA 1.1 draft specification.
Support for new specification features includes:
- Introduction of block scaled tensor types
Block-scaled tensor types represent tensors whose values are grouped into blocks, with each block sharing a scale factor. The dialect supports the OCP microscaling (MX) formats, currently using 32-element blocks along the innermost dimension. The new type is supported by operations including tosa.cast and tosa.const; support across further operations aligned with the EXT-MX-* extensions is ongoing.
- Dynamic shape expression and inference support
Experimental EXT-SHAPE support allows TOSA graphs with unknown dimension sizes to be expressed. Once input shapes are known, shape inference can fold these expressions and propagate static shapes through the graph. In the current specification draft, shape values must resolve to constants during backend compilation; runtime-dynamic shapes are not yet supported.
- Downgrade specification version
A new best-effort transformation rewrites constructs available only in TOSA 1.1.draft into TOSA 1.0-compatible forms where possible. Because not every construct can be downgraded, the resulting IR should subsequently be validated against TOSA 1.0. This helps decouple front-end producers from back-end consumers supporting different TOSA versions.
func.func @main(%arg0: tensor<*xi1>) -> tensor<*xf32> {
%0 = tosa.cast %arg0 : (tensor<*xi1>) -> tensor<*xf32>
return %0 : tensor<*xf32>
}
$ mlir-opt --tosa-downgrade-1-1-to-1-0 test.mlir
func.func @main(%arg0: tensor<*xi1>) -> tensor<*xf32> {
%0 = tosa.cast %arg0 : (tensor<*xi1>) -> tensor<*xi8>
%1 = tosa.cast %0 : (tensor<*xi8>) -> tensor<*xf32>
return %1 : tensor<*xf32>
}
- TOSA to SPIR-V TOSA lowering
A new lowering path that converts TOSA IR to SPIR-V dialect has been added. The lowering produces SPIR-V ARM Graph and Tensor extensions together with the SPIR-V TOSA extended instruction set. Unlike paths through lower-level dialects such as Linalg, Tensor, and Vector, it preserves higher-level TOSA semantics for backends that natively consume SPIR-V Graph and TOSA extensions.
SPIR-V
By Davide Grohmann
LLVM 23 substantially expands machine learning support in MLIR's SPIR-V dialect. The new functionality provides tensor and graph representations, a comprehensive set of TOSA operations, low-precision floating-point types, and an end-to-end lowering path from the TOSA dialect.
- Tensor and graph support
LLVM 23 adds support for the SPV_ARM_tensors and SPV_ARM_graph extensions.
SPV_ARM_tensors introduces SPIR-V tensor types, including ranked and unranked tensors. SPV_ARM_graph represents dataflow computations over these tensors, including graph entry points, inputs, outputs, and constants. The SPIR-V dialect supports parsing, verification, and binary serialization and deserialization for both extensions.
For example, a graph operating on tensors can be represented as:
- TOSA Extended Instruction Set
LLVM 23 adds support for version 001000.1 of the SPIR-V TOSA Extended Instruction Set.
It covers operations for convolution, pooling, matrix multiplication, activation functions, elementwise arithmetic, reductions, data layout, gather and scatter, resize, casting, and rescaling.
These operations preserve TOSA semantics in SPIR-V instead of decomposing them into lower-level arithmetic and control-flow operations. This gives backends that understand TOSA operations more information when compiling an ML workload.
- TOSA-to-SPIR-V lowering
A new lowering path converts the MLIR TOSA dialect directly to the SPIR-V dialect. It combines SPV_ARM_tensors, SPV_ARM_graph, and the TOSA Extended Instruction Set to retain the graph structure, tensor types, and high-level operations of the original program.
For example, the following TOSA function:
can be lowered with: mlir-opt --tosa-to-spirv-tosa input.mlir
The resulting SPIR-V dialect retains both the graph and the ArgMax operation:
The lowering also handles graph constants and automatically derives the SPIR-V extensions and capabilities required by the generated module.
- Float8 support
LLVM 23 adds SPV_EXT_float8 support to the SPIR-V dialect, including the Float8E4M3EXT and Float8E5M2EXT formats. Float8 values can be used in scalar, vector, cooperative-matrix, and Arm tensor types.
The TOSA-to-SPIR-V path can therefore preserve Float8 tensor element types, enabling compact data representation for ML workloads without first promoting values to a wider floating-point type.
- Replicated composite constants
LLVM 23 adds support for SPV_EXT_replicated_composites. This extension represents composite constants whose elements all have the same value without listing every element individually.
The TOSA-to-SPIR-V lowering uses this representation for splat tensor constants, producing more compact SPIR-V modules.
- Graph debugging information
LLVM 23 also adds support for the NonSemantic.Graph.DebugInfo.1 extended instruction set. It represents debugging information for graphs, operations, and tensors without affecting program semantics.
The information is preserved during SPIR-V serialization and deserialization, helping tools relate operations in a generated SPIR-V graph to their compiler representation.
- Experimental ML operations
LLVM 23 supports the Arm.ExperimentalMLOperations.1 extended instruction set. It provides a generic representation for experimental or target-specific ML operations that are not part of the standard TOSA instruction set.
The TOSA-to-SPIR-V lowering can map selected TOSA custom operations to these instructions. This provides an extension point for introducing new operations while their interfaces and semantics are still evolving.
Re-use is only permitted for informational and non-commercial or personal use only.