What is new in LLVM 23? LLVM 23.1.0 was released on August 26, 2026, with almost 1300 patches contributed by Arm teams across architecture support, libraries, tools, performance and ML workloads. The release adds support for the Arm AGI CPU, Arm's first production silicon, based on Armv9.2-A and designed for AI agentic workloads, and native tuning for that CPU improves performance by 1.7% on average on SPEC CPU2017 and 0.8% on average on SPEC CPU2026. LLVM 23 also adds beta ACLE support for Armv9.6-A (SVE2p2, SME2p2) and alpha ACLE support for Armv9.7-A (SVE2p3, SME2p3), plus assembly support for the Extended Hint instruction space (FEAT_HINTE). What is new in LLVM 23? Discover Arm contributions to LLVM 23, including architecture support, performance improvements, code generation, tooling, libraries, and MLIR updates LLVM 23.1.0 https://github.com/llvm/llvm-project/releases release-llvmorg-23.1.0 was released https://discourse.llvm.org/t/llvm-23-1-0-released/91654 on August 26, 2026. Teams across Arm contributed almost 1300 patches to improve architecture support, libraries, tools, performance and increasingly ML workloads. This post provides an overview of the key contributions. To find out more about the previous LLVM release, you can read What is new in LLVM 22? https://developer.arm.com/community/arm-community-blogs/b/tools-software-ides-blog/posts/what-is-new-in-llvm-22 New architecture and CPU support Architecture support By Maciej Gabka LLVM 23 improves architecture support and Arm C Language Extensions ACLE https://support.arm.com/architectures/arm-c-language-extensions support for Arm developers. It also improves architecture correctness in the AArch64 assembler and disassembler, aligning LLVM with the June 2026 updates https://support.arm.com/documentation/109903/2026-06 to the Arm A-profile A64 Instruction Set Architecture . In addition, LLVM 23 adds assembly support for features such as the Extended Hint instruction space FEAT HINTE . LLVM 23 provides support for the beta ACLE specification for data processing features added as part of Armv9.6-A architecture https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2024 , including SVE2p2 and SME2p2, and alpha ACLE specification for Armv9.7-A https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2025 data-processing features, including SVE2p3 and SME2p3. Clang now defines the appropriate ACLE feature-test macros when targeting Armv9.6-A or Armv9.7-A, allowing developers to select architecture-specific code paths at compile time. LLVM 23 improves FP8 ACLE code generation by modeling the FPMR registers as a target-specific memory location. This enables more precise dependency and alias analysis, allowing redundant writes to be eliminated and loop-invariant writes to be safely hoisted out of loops. CPU support LLVM 23 added support for the Arm AGI CPU https://www.arm.com/products/cloud-datacenter/arm-agi-cpu , which is the first production silicon developed by Arm, based on Armv9.2-A and designed specifically for AI agentic workloads. Read more about it in Introducing Arm AGI CPU https://www.arm.com/products/cloud-datacenter/arm-agi-cpu/introduction . LLVM 23 performance benefited from scheduling models for C1-Nano, C1-Premium and C1-Ultra CPUs. Performance improvements Arm AGI CPU improvements By Biplob Mishra LLVM 23 improves tuning for the Arm AGI CPU by adjusting the interleave factor used for vector, scalar, and reduction loops to better match the characteristics of this CPU. These improvements provide measurable gains from native CPU tuning compared with generic tuning. On Arm AGI CPU, native tuning improves performance by 1.7% on average on SPEC CPU2017 and 0.8% on average on SPEC CPU2026. SPEC Improvements By Kiran Chandramohan Arm teams contributed to several improvements in LLVM 23 that boost performance across SPEC benchmarks. In addition to developing new optimizations, the teams continued to monitor SPEC performance throughout the LLVM 23 development cycle and fixed several performance regressions, helping maintain or improve overall performance compared with LLVM 22. Some notable highlights: - 777.zstd r – 4% improvement: Changes to AggressiveInstCombine allow LLVM to recognize split-width 32-bit cttz / ctlz patterns and replace them with wider 64-bit intrinsics, enabling further optimization. - 541.leela r – 3% improvement: Changes to the AArch64 cost model for vector shifts with non-uniform constant shift amounts enable better vectorization decisions. This work originated from the investigation of a regression in imagick and also resulted in a significant improvement for leela . - 753.ns3 r – 1% improvement: Changes to SimplifyCFG exposed additional simplification opportunities by identifying more cases where PHI incoming values lead to undefined behavior. Code generation improvements Extending the capabilities of partial reductions By Sander De Smalen Partial reductions let the vectorizer accumulate into a wider type without first widening every input element. For dot products and similar kernels, this reduces the number of explicit extensions, lowers register pressure, and enables the use of dedicated instructions. LLVM 23 supports many more of these patterns, including floating-point reductions. A half-precision dot product written in ordinary C++ can now make use of SVE fdot when using -ffast-math : js float dot const fp16 a, const fp16 b, int n { float sum = 0.0f; for int i = 0; i < n; ++i sum += float a i float b i ; return sum; } The resulting vector loop uses the more efficient fdot instruction, followed by a horizontal reduction: .Lvector body: ld1h { z1.h }, p0/z, x0, x10, lsl 1 ld1h { z2.h }, p0/z, x1, x10, lsl 1 inch x10 cmp x9, x10 fdot z0.s, z2.h, z1.h b.ne .Lvector body ptrue p0.s faddv s0, p0, z0.s See the full Compiler Explorer example https://godbolt.org/z/MWsr9Er5v . Partial reductions now also work when the operation is predicated inside the loop. Code such as: js int64 t masked dot const int16 t a, const int16 t b, const int16 t mask, int n { int64 t sum = 0; for int i = 0; i < n; ++i if mask i 0 sum += int64 t a i b i ; return sum; } can stay vectorized while ensuring inactive lanes do not contribute to the result. LLVM creates predicates from the mask, uses them for the input loads, and performs the partial reduction with sdot : ld1h { z2.h }, p0/z, x2, x10, lsl 1 cmpgt p1.h, p0/z, z2.h, 0 ld1h { z2.h }, p1/z, x8, x10, lsl 1 ld1h { z3.h }, p1/z, x1, x10, lsl 1 inch x10 sel z2.h, p1, z2.h, z0.h sdot z1.d, z3.h, z2.h The before-and-after output https://godbolt.org/z/v5sf6YTMq shows the full generated function. The matching is no longer limited to a single, canonical sum += product form. LLVM 23 can handle subtraction performed in the middle block: for int i = 0; i < n; ++i sum -= int64 t a i b i ; The vector loop accumulates the products with sdot ; the subtraction from the initial value is then performed in the middle block: .Lvector body: ld1h { z1.h }, p0/z, x0, x10, lsl 1 ld1h { z2.h }, p0/z, x1, x10, lsl 1 inch x10 sdot z0.d, z2.h, z1.h cmp x9, x10 b.ne .Lvector body ptrue p0.d uaddv d0, p0, z0.d fmov x10, d0 sub x2, x2, x10 and mixed add/subtract chains: for int i = 0; i < n; ++i { sum += int64 t a i b i ; sum -= int64 t c i d i ; } In the generated SVE loop, both products use sdot . The subr operations allow the additions and subtractions to be represented in the same partial-reduction chain: ld1h { z1.h }, p0/z, x0, x9, lsl 1 ld1h { z2.h }, p0/z, x1, x9, lsl 1 sdot z0.d, z2.h, z1.h ld1h { z1.h }, p0/z, x2, x9, lsl 1 ld1h { z2.h }, p0/z, x3, x9, lsl 1 subr z0.d, z0.d, 0 sdot z0.d, z2.h, z1.h subr z0.d, z0.d, 0 Explore the generated code for the subtract reduction https://godbolt.org/z/hczYh14hj and the add/sub chain https://godbolt.org/z/fabqqrrq6 . Support has also been added for absolute-difference reductions, a common building block in image, video, and signal-processing code: js unsigned absdiff const unsigned char a, const unsigned char b, int n { unsigned sum = 0; for int i = 0; i < n; ++i sum += builtin abs int a i - int b i ; return sum; } LLVM can now recognize the widening absolute difference as part of a partial reduction instead of materializing a less efficient sequence. In this example, uabd calculates the byte-wise absolute differences and udot with a vector of ones accumulates them into 32-bit lanes: .Lvector body: ldr z2, x12 ldr z3, x13 incb x13 incb x12 subs x14, x14, x8 uabd z2.b, p0/m, z2.b, z3.b udot z0.s, z2.b, z1.b b.ne .Lvector body ptrue p0.s uaddv d0, p0, z0.s fmov w8, s0 See the full Compiler Explorer example https://godbolt.org/z/9dcnodn7r . Vectorizing find-last reductions By Sander De Smalen Conditional scalar assignments inside a loop often encode a “find last” operation: js int last match const int cond, const int values, int n, int needle { int result = -1; for int i = 0; i < n; ++i if cond i == needle result = values i ; return result; } These loops are reductions, even though they do not look like a sum or a minimum. LLVM 23 has better support for recognizing and vectorizing these conditional assignments as find-last reductions. The vector loop compares several elements at once and conditionally retains their indices: .Lvector body: ld1w { z2.s }, p1/z, x0, x8, lsl 2 // load cond i cmpeq p2.s, p1/z, z2.s, z1.s // cond i == needle csetm x11, ne whilelo p3.s, xzr, x11 ld1w { z2.s }, p2/z, x1, x8, lsl 2 // load values i incw x8 mov p0.b, p3/m, p2.b // if any cond i is true, then copy cond to p0 mov z0.s, p3/m, z2.s // if any cond i is true, then copy values to z0 cmp x10, x8 b.ne .Lvector body mov w8, -1 cmp x10, x9 clastb w8, p0, w8, z0.s // extract last active lane, from values This pattern also appears in parsers, scans, and loops that retain the most recent qualifying value. See the full Compiler Explorer example https://godbolt.org/z/TforMxb9d . Explicit floating-point reduction builtins By Sander De Smalen Clang now provides two floating-point reduction builtins: float ordered = builtin reduce in order fadd v, start ; float fast = builtin reduce assoc fadd v, start ; The optional second argument supplies the starting value; without it, the reduction starts from negative zero. The names make the key semantic choice explicit: - builtin reduce in order fadd evaluates the reduction in lane order. - builtin reduce assoc fadd allows reassociation, giving the optimizer freedom to use a tree reduction or target-specific instructions. This allows programmers to specify the required reduction semantics without applying broad fast-math options to the rest of the translation unit. Improved interleaving and SVE shuffles By Sander De Smalen Interleaved data is common in pixels, complex values, structures of channels, and packed sensor input. A lower vectorization factor may be appropriate for a loop, but previously it could prevent LLVM from selecting an interleaved memory operation. LLVM 23 allows interleave shuffles at these lower vectorization factors for SVE, including cases where a tbl sequence efficiently combines deinterleaving with zero-extension. In suitable kernels, one table lookup per output replaces a longer sequence of unpack and shuffle instructions: struct Pixel { unsigned short blue, green, red, alpha; }; void sum channels const Pixel restrict pixels, double restrict sums, unsigned long count { double blue = 0, green = 0, red = 0, alpha = 0; for unsigned long i = 0; i < count; ++i { blue += pixels i .blue; green += pixels i .green; red += pixels i .red; alpha += pixels i .alpha; } sums 0 = blue; sums 1 = green; sums 2 = red; sums 3 = alpha; } Here, converting 16-bit channels to double makes the natural SVE vectorization factor smaller than the four-channel interleave factor. With -O3 -ffast-math -march=armv9-a+sve2 , LLVM vectorizes the loop at vscale x 2 and folds both the deinterleave and zero-extension into the table lookups: tbl z8.h, { z28.h }, z24.h tbl z9.h, { z29.h }, z24.h ucvtf z8.d, p0/m, z8.d ucvtf z9.d, p0/m, z9.d See the full Compiler Explorer example https://godbolt.org/z/e4Pob4aMc . Improved multi-vector loads and stores By Sander De Smalen SVE and SME provide multi-vector memory operations. LLVM 23 improves the use of the instruction's addressing modes. For example, js include