# What is new in LLVM 23?

> Source: <https://developer.arm.com/community/arm-community-blogs/b/tools-software-ides-blog/posts/what-is-new-in-llvm-23>
> Published: 2026-10-03 20:03:30+00:00

# What is new in LLVM 23?

Discover Arm contributions to LLVM 23, including architecture support, performance improvements, code generation, tooling, libraries, and MLIR updates

[LLVM 23.1.0](https://github.com/llvm/llvm-project/releases#release-llvmorg-23.1.0) was [released](https://discourse.llvm.org/t/llvm-23-1-0-released/91654) on August 26, 2026. Teams across Arm contributed almost 1300 patches to improve architecture support, libraries, tools, performance and increasingly ML workloads. This post provides an overview of the key contributions.

To find out more about the previous LLVM release, you can read [What is new in LLVM 22?](https://developer.arm.com/community/arm-community-blogs/b/tools-software-ides-blog/posts/what-is-new-in-llvm-22)

## New architecture and CPU support

### Architecture support

By Maciej Gabka

LLVM 23 improves architecture support and [Arm C Language Extensions (ACLE)](https://support.arm.com/architectures/arm-c-language-extensions) support for Arm developers. It also improves architecture correctness in the AArch64 assembler and disassembler, aligning LLVM with the [June 2026 updates](https://support.arm.com/documentation/109903/2026-06) to the Arm A-profile A64 Instruction Set Architecture**.** In addition, LLVM 23 adds assembly support for features such as the Extended Hint instruction space (`FEAT_HINTE`).

LLVM 23 provides support for the beta ACLE specification for data processing features added as part of [Armv9.6-A architecture](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2024), including SVE2p2 and SME2p2, and alpha ACLE specification for [Armv9.7-A](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2025) data-processing features, including SVE2p3 and SME2p3. Clang now defines the appropriate ACLE feature-test macros when targeting Armv9.6-A or Armv9.7-A, allowing developers to select architecture-specific code paths at compile time.

LLVM 23 improves FP8 ACLE code generation by modeling the FPMR registers as a target-specific memory location. This enables more precise dependency and alias analysis, allowing redundant writes to be eliminated and loop-invariant writes to be safely hoisted out of loops.

### CPU support

LLVM 23 added support for the [Arm AGI CPU](https://www.arm.com/products/cloud-datacenter/arm-agi-cpu), which is the first production silicon developed by Arm, based on Armv9.2-A and designed specifically for AI agentic workloads. Read more about it in [Introducing Arm AGI CPU](https://www.arm.com/products/cloud-datacenter/arm-agi-cpu/introduction).

LLVM 23 performance benefited from scheduling models for C1-Nano, C1-Premium and C1-Ultra CPUs.

## Performance improvements

### Arm AGI CPU improvements

By Biplob Mishra

LLVM 23 improves tuning for the Arm AGI CPU by adjusting the interleave factor used for vector, scalar, and reduction loops to better match the characteristics of this CPU. These improvements provide measurable gains from native CPU tuning compared with generic tuning. On Arm AGI CPU, native tuning improves performance by 1.7% on average on SPEC CPU2017 and 0.8% on average on SPEC CPU2026.

### SPEC Improvements

By Kiran Chandramohan

Arm teams contributed to several improvements in LLVM 23 that boost performance across SPEC benchmarks. In addition to developing new optimizations, the teams continued to monitor SPEC performance throughout the LLVM 23 development cycle and fixed several performance regressions, helping maintain or improve overall performance compared with LLVM 22.

Some notable highlights:

- 777.zstd_r – 4% improvement: Changes to `AggressiveInstCombine` allow LLVM to recognize split-width 32-bit`cttz` /`ctlz` patterns and replace them with wider 64-bit intrinsics, enabling further optimization.
- 541.leela_r – 3% improvement: Changes to the AArch64 cost model for vector shifts with non-uniform constant shift amounts enable better vectorization decisions. This work originated from the investigation of a regression in `imagick` and also resulted in a significant improvement for`leela` .
- 753.ns3_r – 1% improvement: Changes to `SimplifyCFG` exposed additional simplification opportunities by identifying more cases where PHI incoming values lead to undefined behavior.

## Code generation improvements

### Extending the capabilities of partial reductions

By Sander De Smalen

Partial reductions let the vectorizer accumulate into a wider type without first widening every input element. For dot products and similar kernels, this reduces the number of explicit extensions, lowers register pressure, and enables the use of dedicated instructions.

LLVM 23 supports many more of these patterns, including floating-point reductions. A half-precision dot product written in ordinary C++ can now make use of SVE `fdot` when using `-ffast-math`:

``` js
float dot(const __fp16 *a, const __fp16 *b, int n) {
  float sum = 0.0f;
  for (int i = 0; i < n; ++i)
    sum += (float)a[i] * (float)b[i];
  return sum;
}
```

The resulting vector loop uses the more efficient `fdot` instruction, followed by a horizontal reduction:

```
.Lvector_body:
    ld1h    { z1.h }, p0/z, [x0, x10, lsl #1]
    ld1h    { z2.h }, p0/z, [x1, x10, lsl #1]
    inch    x10
    cmp     x9, x10
    fdot    z0.s, z2.h, z1.h
    b.ne    .Lvector_body

    ptrue   p0.s
    faddv   s0, p0, z0.s
```

See the full [Compiler Explorer example](https://godbolt.org/z/MWsr9Er5v).

Partial reductions now also work when the operation is predicated inside the loop. Code such as:

``` js
int64_t masked_dot(const int16_t *a, const int16_t *b,
                   const int16_t *mask, int n) {
  int64_t sum = 0;
  for (int i = 0; i < n; ++i)
    if (mask[i] > 0)
      sum += (int64_t)a[i] * b[i];
  return sum;
}
```

can stay vectorized while ensuring inactive lanes do not contribute to the result. LLVM creates predicates from the mask, uses them for the input loads, and performs the partial reduction with `sdot`:

```
    ld1h    { z2.h }, p0/z, [x2, x10, lsl #1]
    cmpgt   p1.h, p0/z, z2.h, #0
    ld1h    { z2.h }, p1/z, [x8, x10, lsl #1]
    ld1h    { z3.h }, p1/z, [x1, x10, lsl #1]
    inch    x10
    sel     z2.h, p1, z2.h, z0.h
    sdot    z1.d, z3.h, z2.h
```

The [before-and-after output](https://godbolt.org/z/v5sf6YTMq) shows the full generated function.

The matching is no longer limited to a single, canonical `sum += product` form. LLVM 23 can handle subtraction performed in the middle block:

```
for (int i = 0; i < n; ++i)
  sum -= (int64_t)a[i] * b[i];
```

The vector loop accumulates the products with `sdot`; the subtraction from the initial value is then performed in the middle block:

```
.Lvector_body:
    ld1h    { z1.h }, p0/z, [x0, x10, lsl #1]
    ld1h    { z2.h }, p0/z, [x1, x10, lsl #1]
    inch    x10
    sdot    z0.d, z2.h, z1.h
    cmp     x9, x10
    b.ne    .Lvector_body

    ptrue   p0.d
    uaddv   d0, p0, z0.d
    fmov    x10, d0
    sub     x2, x2, x10
```

and mixed add/subtract chains:

```
for (int i = 0; i < n; ++i) {
  sum += (int64_t)a[i] * b[i];
  sum -= (int64_t)c[i] * d[i];
}
```

In the generated SVE loop, both products use `sdot`. The `subr` operations allow the additions and subtractions to be represented in the same partial-reduction chain:

```
    ld1h    { z1.h }, p0/z, [x0, x9, lsl #1]
    ld1h    { z2.h }, p0/z, [x1, x9, lsl #1]
    sdot    z0.d, z2.h, z1.h
    ld1h    { z1.h }, p0/z, [x2, x9, lsl #1]
    ld1h    { z2.h }, p0/z, [x3, x9, lsl #1]
    subr    z0.d, z0.d, #0
    sdot    z0.d, z2.h, z1.h
    subr    z0.d, z0.d, #0
```

Explore the generated code for the [subtract reduction](https://godbolt.org/z/hczYh14hj) and the [add/sub chain](https://godbolt.org/z/fabqqrrq6).

Support has also been added for absolute-difference reductions, a common building block in image, video, and signal-processing code:

``` js
unsigned absdiff(const unsigned char *a, const unsigned char *b, int n) {
  unsigned sum = 0;
  for (int i = 0; i < n; ++i)
    sum += __builtin_abs((int)a[i] - (int)b[i]);
  return sum;
}
```

LLVM can now recognize the widening absolute difference as part of a partial reduction instead of materializing a less efficient sequence. In this example, `uabd` calculates the byte-wise absolute differences and `udot` with a vector of ones accumulates them into 32-bit lanes:

```
.Lvector_body:
    ldr     z2, [x12]
    ldr     z3, [x13]
    incb    x13
    incb    x12
    subs    x14, x14, x8
    uabd    z2.b, p0/m, z2.b, z3.b
    udot    z0.s, z2.b, z1.b
    b.ne    .Lvector_body

    ptrue   p0.s
    uaddv   d0, p0, z0.s
    fmov    w8, s0
```

See the full [Compiler Explorer example](https://godbolt.org/z/9dcnodn7r).

### Vectorizing find-last reductions

By Sander De Smalen

Conditional scalar assignments inside a loop often encode a “find last” operation:

``` js
int last_match(const int *cond, const int *values, int n, int needle) {
  int result = -1;
  for (int i = 0; i < n; ++i)
    if (cond[i] == needle)
      result = values[i];
  return result;
}
```

These loops are reductions, even though they do not look like a sum or a minimum. LLVM 23 has better support for recognizing and vectorizing these conditional assignments as find-last reductions. The vector loop compares several elements at once and conditionally retains their indices:

```
.Lvector_body:
    ld1w    { z2.s }, p1/z, [x0, x8, lsl #2]    // load cond[i]
    cmpeq   p2.s, p1/z, z2.s, z1.s              // cond[i] == needle
    csetm   x11, ne
    whilelo p3.s, xzr, x11
    ld1w    { z2.s }, p2/z, [x1, x8, lsl #2]    // load values[i]
    incw    x8
    mov     p0.b, p3/m, p2.b                    // if any cond[i] is true, then copy cond[] to p0
    mov     z0.s, p3/m, z2.s                    // if any cond[i] is true, then copy values[] to z0
    cmp     x10, x8
    b.ne    .Lvector_body

    mov     w8, #-1
    cmp     x10, x9
    clastb  w8, p0, w8, z0.s                    // extract last active lane, from values[]
```

This pattern also appears in parsers, scans, and loops that retain the most recent qualifying value.

See the full [Compiler Explorer example](https://godbolt.org/z/TforMxb9d).

### Explicit floating-point reduction builtins

By Sander De Smalen

Clang now provides two floating-point reduction builtins:

```
float ordered = __builtin_reduce_in_order_fadd(v, start);
float fast    = __builtin_reduce_assoc_fadd(v, start);
```

The optional second argument supplies the starting value; without it, the reduction starts from negative zero. The names make the key semantic choice explicit:

- `__builtin_reduce_in_order_fadd` evaluates the reduction in lane order.
- `__builtin_reduce_assoc_fadd` allows reassociation, giving the optimizer freedom to use a tree reduction or target-specific instructions.

This allows programmers to specify the required reduction semantics without applying broad fast-math options to the rest of the translation unit.

### Improved interleaving and SVE shuffles

By Sander De Smalen

Interleaved data is common in pixels, complex values, structures of channels, and packed sensor input. A lower vectorization factor may be appropriate for a loop, but previously it could prevent LLVM from selecting an interleaved memory operation.

LLVM 23 allows interleave shuffles at these lower vectorization factors for SVE, including cases where a `tbl` sequence efficiently combines deinterleaving with zero-extension. In suitable kernels, one table lookup per output replaces a longer sequence of unpack and shuffle instructions:

```
struct Pixel { unsigned short blue, green, red, alpha; };

void sum_channels(const Pixel *__restrict pixels,
                  double *__restrict sums, unsigned long count) {
  double blue = 0, green = 0, red = 0, alpha = 0;
  for (unsigned long i = 0; i < count; ++i) {
    blue  += pixels[i].blue;
    green += pixels[i].green;
    red   += pixels[i].red;
    alpha += pixels[i].alpha;
  }
  sums[0] = blue; sums[1] = green;
  sums[2] = red;  sums[3] = alpha;
}
```

Here, converting 16-bit channels to `double` makes the natural SVE vectorization factor smaller than the four-channel interleave factor. With `-O3 -ffast-math -march=armv9-a+sve2`, LLVM vectorizes the loop at `vscale x 2` and folds both the deinterleave and zero-extension into the table lookups:

```
tbl     z8.h, { z28.h }, z24.h
tbl     z9.h, { z29.h }, z24.h
ucvtf   z8.d, p0/m, z8.d
ucvtf   z9.d, p0/m, z9.d
```

See the full [Compiler Explorer example](https://godbolt.org/z/e4Pob4aMc).

### Improved multi-vector loads and stores

By Sander De Smalen

SVE and SME provide multi-vector memory operations. LLVM 23 improves the use of the instruction's addressing modes. For example,

``` js
#include <arm_sve.h>

void copy_vector_pair(unsigned char *dst, const unsigned char *src) {
  svcount_t all = svptrue_c8();
  svuint8x2_t data = svld1_vnum_u8_x2(all, src, -2);
  svst1_vnum_u8_x2(all, dst, 14, data);
}
```

LLVM 23 folds both offsets directly into the multi-vector instructions, avoiding separate pointer adjustments:

```
ptrue   pn8.b
ld1b    { z0.b, z1.b }, pn8/z, [x1, #-2, mul vl]
st1b    { z0.b, z1.b }, pn8,   [x0, #14, mul vl]
```

See the full [Compiler Explorer example](https://godbolt.org/z/r6v6fYnGf)

Streaming-mode for SME2 code also benefits from enabling sub-register liveness. For example, this kernel loads two rows as vector pairs, then transposes the tuple view for a dot product:

``` js
#include <arm_sme.h>

void dot_two_rows(unsigned long stride, const unsigned char *ptr)
    __arm_streaming __arm_inout("za") {
  svcount_t all = svptrue_c8();
  svuint8x2_t row0 = svld1_u8_x2(all, ptr);
  svuint8x2_t row1 = svld1_u8_x2(all, ptr + stride);
  svuint8x2_t lhs = svcreate2_u8(svget2_u8(row0, 0), svget2_u8(row1, 0));
  svuint8x2_t rhs = svcreate2_u8(svget2_u8(row0, 1), svget2_u8(row1, 1));
  svdot_za32_u8_vg1x2(0, lhs, rhs);
}
```

LLVM 22 selects strided multi-vector loads, but must copy half of each tuple into consecutive registers for `udot`:

```
ld1b    { z16.b, z24.b }, pn8/z, [x1]
ld1b    { z17.b, z25.b }, pn8/z, [x1, x0]
mov     z0.d, z24.d
mov     z1.d, z25.d
udot    za.s[w8, 0, vgx2], { z16.b, z17.b }, { z0.b, z1.b }
```

LLVM 23 can keep both tuple views live and removes the copies:

```
ld1b    { z16.b, z24.b }, pn8/z, [x1]
ld1b    { z17.b, z25.b }, pn8/z, [x1, x0]
udot    za.s[w8, 0, vgx2], { z16.b, z17.b }, { z24.b, z25.b }
```

See the full [Compiler Explorer example](https://godbolt.org/z/aacj3jocW)

### Improved use of SVE immediates

By Sander De Smalen

LLVM 23 recognizes more opportunities to use SVE instructions that have an immediate form, instead of first materializing a constant in a vector register and using an equivalent NEON instruction that lacks an immediate form. That saves an instruction and, just as importantly, one temporary register.

For example:

```
#include <arm_neon.h>

uint32x4_t increment(uint32x4_t x) {
  return vaddq_u32(x, vdupq_n_u32(1));
}
```

LLVM 22 materialized the splat in a second vector register before using a NEON add:

```
movi    v1.4s, #1
add     v0.4s, v0.4s, v1.4s
```

LLVM 23 can instead use the SVE immediate form directly:

```
add     z0.s, z0.s, #1
```

See the [Compiler Explorer example](https://godbolt.org/z/x5a5138vv)

### Additional instruction-selection improvements

By Sander De Smalen

LLVM 23 also includes local AArch64 combines that remove redundant instructions.

One example folds a condition-code materialization followed by a bit-test branch:

```
cset    w8, eq
tbnz    w8, #0, .Lmatch
```

directly into the equivalent conditional branch:

```
b.eq    .Lmatch
```

Although individually small, these folds can reduce code size and leave fewer instructions for later scheduling stages.

### Improved carry-less multiplication

By Matthew Devereau

LLVM 23 improves AArch64 code generation for carry-less multiplication, an operation commonly used in cryptography and data-integrity algorithms. More fixed-width and scalable-vector forms are now mapped to sequences that utilize `PMUL`/` PMULL` instructions when SVE2/SME and AES extensions are available and scalar forms can now utilize the vector register. The cost model has also been updated so that LLVM can make better optimization decisions. 

### Faster compilation with GlobalISel

By Cullen Rhodes

Compile-time has improved significantly in LLVM 23. Debug builds (`-O0 -g`) are 15% faster on CTMark, with some workloads such as SQLite almost 25% faster. Optimized builds (`-O3`) are also 8.7% faster. This is largely due to extensive improvements in `GlobalISel`, the default instruction selector at `-O0`, as well as broader improvements in Clang debug info and LLVM code generation.

## Tools improvements

### BOLT improvements

By Paschalis Mpeis

LLVM 23 improves BOLT support for AArch64 binaries, with contributions focused on Branch Target Identification (BTI), compare-and-branch instructions, long-jump code layout, and usability.

We introduced initial BOLT support for processing BTI-enabled AArch64 binaries. BOLT can now check and patch PLT entries and indirect branch targets where possible. It can disassemble PLT entries in BTI binaries, patch LLD-generated PLTs with BTI landing pads, and patch ignored functions in place when indirect branches target them.

We also added compact-code-model support for Armv9.6-A `FEAT_CMPBR` compare-and-branch instructions. This includes support for block reordering, function splitting, branch inversion where legal, and trampoline handling for difficult branch cases.

This release also improves AArch64 usability. Unsupported or target-specific BOLT passes now report clear errors instead of crashing or silently doing nothing, and new documentation records AArch64 optimization flag support. BOLT also gains support for `SHT_CREL` code relocations, reports compressed debug sections cleanly, improves `--hugify` long-jump layout handling, and expands Arm test coverage.

### Memory Protection Keys and Permission Overlays support in LLDB

By David Spickett

Memory permissions are typically set in the page tables. Programs often change these permissions at runtime. For example they might initialise key structures and then make them read-only to protect them against attackers trying to modify them later.

However, changing page table permissions is expensive. On Linux it is usually done with the `mprotect` system call. Linux has an alternative called ["Memory Protection Keys"](https://docs.kernel.org/core-api/protection-keys.html) which allows permissions to be changed without a system call. These keys work together with a hardware feature which stores "permission overlays". On AArch64 this is the Permission Overlay Extension (`FEAT_S1POE)` introduced in Armv8.8.

New in LLDB 23 is support for debugging Memory Protection Keys and permission overlays on AArch64 Linux.

Below is an example of a protection key fault:

```
(lldb) c
Process 462 resuming
Process 462 stopped
<...> stop reason = signal SIGSEGV: failed protection key checks (fault address=0xffffff7d60000)
<...>
-> 106       read_only_page[0] = '?';
```

This program tried to write to memory and was not allowed to do so. You can tell that the page table entry did have write permissions because of the description of the `SIGSEGV` signal. It refers specifically to protection keys (`SEGV_PKUERR)` , as opposed to "invalid permissions for mapped object" (`SEGV_ACCERR)` which is generated when the page table permissions are the cause.

You can check the permissions with the `memory region` command:

```
(lldb) memory region read_only_page
[0x000ffffff7d60000-0x000ffffff7d70000) rw-
protection key: 6 (r--, effective: r--)
```

In prior versions of LLDB, all you would see is the `rw-` on the 2nd line. This represents the permissions in the page table. They allow reading (`r`) and writing (` w`), but not execution (` x`).

New in LLDB 23 is the 3rd line. It shows:

- The protection key assigned to this memory mapping (`6` ).
- The permission overlay that the key refers to (`r--)` .
- The effective permissions (`r--)` after combining the page table permissions with the permission overlay.

Permission overlays can only remove permissions. In this example `rw-` was overlaid with `r--` . This overlay keeps the `r` permission but removes the `w` and `x` permissions. That is why the program failed to write. This memory mapping is read-only.

LLDB also lets you access the `POR_EL0` register where the permission overlays are stored. To read about this and other AArch64 Linux features, see the [LLDB documentation](https://lldb.llvm.org/use/aarch64-linux.html).

### Flang improvements

By Tom Eccles

Arm's contributions to LLVM 23 improved Flang's compatibility, performance and OpenMP support. Taskloop lowering is now more robust, with better handling of loop bounds, steps and privatized character values. Flang also supports the non-standard RTC intrinsic used by some existing Fortran applications. A new opt-in mode (`-freal-sum-association` in LLVM 23 or `-ffp-sum-association` in LLVM 24) can reassociate real-valued sums where permitted by the Fortran standard, exposing additional optimization opportunities while preserving explicit parentheses. Our biggest contribution was continuing upstream maintenance for flang's internal MLIR dialects, OpenMP on CPU support, and Windows support. I was the most active reviewer for flang OpenMP support (approximately 18% of reviews), and ranked third in flang as a whole (approximately 9% of reviews).

### BOLT optimization for Flang

By Pawel Osmialowski

The purpose of BOLTing (optimizing with BOLT) the compiler itself is to make it compile the code faster; this is not about the optimization capabilities of the compiler itself (a common misconception).

Following the already existing methodology for having the Clang compiler BOLTed, we have contributed a similar method for BOLTing the Flang compiler. As with Clang, our approach is also based on the CMake caches and reuses the in-tree profiling utilities that were already used for BOLTing Clang. The proposed CMake caches introduce two optimization methods: BOLT and BOLT+PGO (with and without LTO). We did our performance evaluation of the BOLT+PGO method on different servers using three Fortran codebase examples (the DBCSR library, the Polyhedron benchmark, and the LAPACK library) and saw the following speedup in compilation times:

- DBCSR: 12.9 – 21.9%
- Polyhedron (pb11): 12.7 – 18.7%
- LAPACK-3.6.1: 15.3 – 16.2%

During this work, however, we also discovered that the PGO-based optimization process is affected by a race condition that can cause the performance-training stage to hang intermittently. Until this issue is resolved, we recommend using BOLT without PGO when optimizing Flang.

### Windows on Arm and Arm64EC support

By David Truby

LLVM 23 extends Windows on Arm support across OpenMP, Flang and native interoperability. The OpenMP runtime now supports Arm64EC and Arm64X configurations, making it easier to combine Arm-native and emulated code. Flang’s runtime gains improved Windows implementations for date and time queries and file handling, together with Arm64EC build fixes. LLVM can also generate Arm64EC interoperability thunks for functions using `bfloat16` values.

## Libraries improvements

### Optimized AArch32 software floating point functions

By Simon Tatham

In the `compiler-rt` builtins library, there is now a full set of highly optimized implementations of floating-point basic arithmetic, for AArch32 systems supporting either of the Arm or Thumb2 instruction set but with no hardware FPU. A small number of functions also have optimized Thumb1 implementations, suitable for Armv6-M or Armv8-M Baseline. On average, the new functions run roughly twice as fast as the generic functions they replaced.

### Libc vector math

By Dylan Fleming

LLVM 23 introduces a `mathvec` component into the LLVM libc, providing LLVM vector math implementations of standard math functions that can ship alongside their scalar counterparts. Current work has focused on establishing the component structure, testing infrastructure, and delivering proofs of concept for both a cross-architecture (generic) and an AdvSIMD-optimised single-precision exponential (`expf`).

Like scalar routines, vector routines are correctly rounded: they return the floating-point value nearest to the exact mathematical result for the supported rounding mode (presently, round to nearest, with ties to even). While this level of accuracy may be excessive for many use cases, correct rounding offers additional benefits.

Correctly rounded functions preserve properties of mathematical functions that lower-accuracy implementations may lose, potentially leading to unintuitive outputs, logical errors, and poor numerical stability. Among the properties frequently required by applications, monotonicity is preserved. For example, `x <= y` implies `f(x) <= f(y)` holds for floating-point quantities too. Perhaps more noticeably, functions remain within their true bounds, such as `|sin(x)| <= 1` . Correct rounding also guarantees bitwise-identical results across:

- Scalar and vector implementations, enabling safe auto-vectorization without changing numerical behavior.
- Architectures and platforms, improving reproducibility and avoiding target-dependent production results.
- Library versions, helping users adopt the latest version and optimizations seamlessly.

Going forward, we plan to increase the coverage of vector routines, initially focusing on FP32 and FP16 before expanding to FP64 and BF16. These will first be delivered in generic form and enabled on all supported architectures, while target-specific optimizations will follow where observable gains can be achieved.

## MLIR Improvements

### TOSA

By Luke Hutton

The TOSA MLIR dialect in LLVM 23 continues to support TOSA 1.0 while introducing early support for new features from the TOSA 1.1 draft [specification](https://github.com/arm/tosa-specification).

Support for new specification features includes:

- Introduction of block scaled tensor types

Block-scaled tensor types represent tensors whose values are grouped into blocks, with each block sharing a scale factor. The dialect supports the [OCP microscaling](https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf) (MX) formats, currently using 32-element blocks along the innermost dimension. The new type is supported by operations including `tosa.cast` and `tosa.const`; support across further operations aligned with the `EXT-MX-*` extensions is ongoing.

- Dynamic shape expression and inference support

Experimental `EXT-SHAPE` support allows TOSA graphs with unknown dimension sizes to be expressed. Once input shapes are known, shape inference can fold these expressions and propagate static shapes through the graph. In the current specification draft, shape values must resolve to constants during backend compilation; runtime-dynamic shapes are not yet supported.

- Downgrade specification version

A new best-effort transformation rewrites constructs available only in TOSA 1.1.draft into TOSA 1.0-compatible forms where possible. Because not every construct can be downgraded, the resulting IR should subsequently be validated against TOSA 1.0. This helps decouple front-end producers from back-end consumers supporting different TOSA versions.

``` php
func.func @main(%arg0: tensor<*xi1>) -> tensor<*xf32> {
  %0 = tosa.cast %arg0 : (tensor<*xi1>) -> tensor<*xf32>
  return %0 : tensor<*xf32>
}

$ mlir-opt --tosa-downgrade-1-1-to-1-0 test.mlir

func.func @main(%arg0: tensor<*xi1>) -> tensor<*xf32> {
  %0 = tosa.cast %arg0 : (tensor<*xi1>) -> tensor<*xi8>
  %1 = tosa.cast %0 : (tensor<*xi8>) -> tensor<*xf32>
  return %1 : tensor<*xf32>
}
```

- TOSA to SPIR-V TOSA lowering

A new lowering path that converts TOSA IR to SPIR-V dialect has been added. The lowering produces SPIR-V ARM Graph and Tensor extensions together with the SPIR-V TOSA extended instruction set. Unlike paths through lower-level dialects such as Linalg, Tensor, and Vector, it preserves higher-level TOSA semantics for backends that natively consume SPIR-V Graph and TOSA extensions.

### SPIR-V

By Davide Grohmann

LLVM 23 substantially expands machine learning support in MLIR's SPIR-V dialect. The new functionality provides tensor and graph representations, a comprehensive set of TOSA operations, low-precision floating-point types, and an end-to-end lowering path from the TOSA dialect.

- Tensor and graph support

LLVM 23 adds support for the [SPV_ARM_tensors](https://github.khronos.org/SPIRV-Registry/extensions/ARM/SPV_ARM_tensors.html) and [SPV_ARM_graph](https://github.khronos.org/SPIRV-Registry/extensions/ARM/SPV_ARM_graph.html) extensions.

SPV_ARM_tensors introduces SPIR-V tensor types, including ranked and unranked tensors. SPV_ARM_graph represents dataflow computations over these tensors, including graph entry points, inputs, outputs, and constants. The SPIR-V dialect supports parsing, verification, and binary serialization and deserialization for both extensions.

For example, a graph operating on tensors can be represented as:

- TOSA Extended Instruction Set

LLVM 23 adds support for version 001000.1 of the [SPIR-V TOSA Extended Instruction Set](https://github.khronos.org/SPIRV-Registry/extended/TOSA.001000.1.html).

It covers operations for convolution, pooling, matrix multiplication, activation functions, elementwise arithmetic, reductions, data layout, gather and scatter, resize, casting, and rescaling.

These operations preserve TOSA semantics in SPIR-V instead of decomposing them into lower-level arithmetic and control-flow operations. This gives backends that understand TOSA operations more information when compiling an ML workload.

- TOSA-to-SPIR-V lowering

A new lowering path converts the MLIR TOSA dialect directly to the SPIR-V dialect. It combines SPV_ARM_tensors, SPV_ARM_graph, and the TOSA Extended Instruction Set to retain the graph structure, tensor types, and high-level operations of the original program.

For example, the following TOSA function:

can be lowered with: `mlir-opt --tosa-to-spirv-tosa input.mlir`

The resulting SPIR-V dialect retains both the graph and the ArgMax operation:

The lowering also handles graph constants and automatically derives the SPIR-V extensions and capabilities required by the generated module.

- Float8 support

LLVM 23 adds [SPV_EXT_float8](https://github.khronos.org/SPIRV-Registry/extensions/EXT/SPV_EXT_float8.html) support to the SPIR-V dialect, including the Float8E4M3EXT and Float8E5M2EXT formats. Float8 values can be used in scalar, vector, cooperative-matrix, and Arm tensor types.

The TOSA-to-SPIR-V path can therefore preserve Float8 tensor element types, enabling compact data representation for ML workloads without first promoting values to a wider floating-point type.

- Replicated composite constants

LLVM 23 adds support for [SPV_EXT_replicated_composites](https://github.khronos.org/SPIRV-Registry/extensions/EXT/SPV_EXT_replicated_composites.html). This extension represents composite constants whose elements all have the same value without listing every element individually.

The TOSA-to-SPIR-V lowering uses this representation for splat tensor constants, producing more compact SPIR-V modules.

- Graph debugging information

LLVM 23 also adds support for the [NonSemantic.Graph.DebugInfo.1](https://github.khronos.org/SPIRV-Registry/nonsemantic/NonSemantic.Graph.DebugInfo.html) extended instruction set. It represents debugging information for graphs, operations, and tensors without affecting program semantics.

The information is preserved during SPIR-V serialization and deserialization, helping tools relate operations in a generated SPIR-V graph to their compiler representation.

- Experimental ML operations

LLVM 23 supports the [Arm.ExperimentalMLOperations](https://github.khronos.org/SPIRV-Registry/extended/Arm.ExperimentalMLOperations.html).1 extended instruction set. It provides a generic representation for experimental or target-specific ML operations that are not part of the standard TOSA instruction set.

The TOSA-to-SPIR-V lowering can map selected TOSA custom operations to these instructions. This provides an extension point for introducing new operations while their interfaces and semantics are still evolving.

Re-use is only permitted for informational and non-commercial or personal use only.
