{"slug": "what-is-new-in-llvm-23", "title": "What is new in LLVM 23?", "summary": "LLVM 23.1.0 was released on August 26, 2026, with almost 1300 patches contributed by Arm teams across architecture support, libraries, tools, performance and ML workloads. The release adds support for the Arm AGI CPU, Arm's first production silicon, based on Armv9.2-A and designed for AI agentic workloads, and native tuning for that CPU improves performance by 1.7% on average on SPEC CPU2017 and 0.8% on average on SPEC CPU2026. LLVM 23 also adds beta ACLE support for Armv9.6-A (SVE2p2, SME2p2) and alpha ACLE support for Armv9.7-A (SVE2p3, SME2p3), plus assembly support for the Extended Hint instruction space (FEAT_HINTE).", "body_md": "# What is new in LLVM 23?\n\nDiscover Arm contributions to LLVM 23, including architecture support, performance improvements, code generation, tooling, libraries, and MLIR updates\n\n[LLVM 23.1.0](https://github.com/llvm/llvm-project/releases#release-llvmorg-23.1.0) was [released](https://discourse.llvm.org/t/llvm-23-1-0-released/91654) on August 26, 2026. Teams across Arm contributed almost 1300 patches to improve architecture support, libraries, tools, performance and increasingly ML workloads. This post provides an overview of the key contributions.\n\nTo find out more about the previous LLVM release, you can read [What is new in LLVM 22?](https://developer.arm.com/community/arm-community-blogs/b/tools-software-ides-blog/posts/what-is-new-in-llvm-22)\n\n## New architecture and CPU support\n\n### Architecture support\n\nBy Maciej Gabka\n\nLLVM 23 improves architecture support and [Arm C Language Extensions (ACLE)](https://support.arm.com/architectures/arm-c-language-extensions) support for Arm developers. It also improves architecture correctness in the AArch64 assembler and disassembler, aligning LLVM with the [June 2026 updates](https://support.arm.com/documentation/109903/2026-06) to the Arm A-profile A64 Instruction Set Architecture**.** In addition, LLVM 23 adds assembly support for features such as the Extended Hint instruction space (`FEAT_HINTE`).\n\nLLVM 23 provides support for the beta ACLE specification for data processing features added as part of [Armv9.6-A architecture](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2024), including SVE2p2 and SME2p2, and alpha ACLE specification for [Armv9.7-A](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-a-profile-architecture-developments-2025) data-processing features, including SVE2p3 and SME2p3. Clang now defines the appropriate ACLE feature-test macros when targeting Armv9.6-A or Armv9.7-A, allowing developers to select architecture-specific code paths at compile time.\n\nLLVM 23 improves FP8 ACLE code generation by modeling the FPMR registers as a target-specific memory location. This enables more precise dependency and alias analysis, allowing redundant writes to be eliminated and loop-invariant writes to be safely hoisted out of loops.\n\n### CPU support\n\nLLVM 23 added support for the [Arm AGI CPU](https://www.arm.com/products/cloud-datacenter/arm-agi-cpu), which is the first production silicon developed by Arm, based on Armv9.2-A and designed specifically for AI agentic workloads. Read more about it in [Introducing Arm AGI CPU](https://www.arm.com/products/cloud-datacenter/arm-agi-cpu/introduction).\n\nLLVM 23 performance benefited from scheduling models for C1-Nano, C1-Premium and C1-Ultra CPUs.\n\n## Performance improvements\n\n### Arm AGI CPU improvements\n\nBy Biplob Mishra\n\nLLVM 23 improves tuning for the Arm AGI CPU by adjusting the interleave factor used for vector, scalar, and reduction loops to better match the characteristics of this CPU. These improvements provide measurable gains from native CPU tuning compared with generic tuning. On Arm AGI CPU, native tuning improves performance by 1.7% on average on SPEC CPU2017 and 0.8% on average on SPEC CPU2026.\n\n### SPEC Improvements\n\nBy Kiran Chandramohan\n\nArm teams contributed to several improvements in LLVM 23 that boost performance across SPEC benchmarks. In addition to developing new optimizations, the teams continued to monitor SPEC performance throughout the LLVM 23 development cycle and fixed several performance regressions, helping maintain or improve overall performance compared with LLVM 22.\n\nSome notable highlights:\n\n- 777.zstd_r – 4% improvement: Changes to `AggressiveInstCombine` allow LLVM to recognize split-width 32-bit`cttz` /`ctlz` patterns and replace them with wider 64-bit intrinsics, enabling further optimization.\n- 541.leela_r – 3% improvement: Changes to the AArch64 cost model for vector shifts with non-uniform constant shift amounts enable better vectorization decisions. This work originated from the investigation of a regression in `imagick` and also resulted in a significant improvement for`leela` .\n- 753.ns3_r – 1% improvement: Changes to `SimplifyCFG` exposed additional simplification opportunities by identifying more cases where PHI incoming values lead to undefined behavior.\n\n## Code generation improvements\n\n### Extending the capabilities of partial reductions\n\nBy Sander De Smalen\n\nPartial reductions let the vectorizer accumulate into a wider type without first widening every input element. For dot products and similar kernels, this reduces the number of explicit extensions, lowers register pressure, and enables the use of dedicated instructions.\n\nLLVM 23 supports many more of these patterns, including floating-point reductions. A half-precision dot product written in ordinary C++ can now make use of SVE `fdot` when using `-ffast-math`:\n\n``` js\nfloat dot(const __fp16 *a, const __fp16 *b, int n) {\n  float sum = 0.0f;\n  for (int i = 0; i < n; ++i)\n    sum += (float)a[i] * (float)b[i];\n  return sum;\n}\n```\n\nThe resulting vector loop uses the more efficient `fdot` instruction, followed by a horizontal reduction:\n\n```\n.Lvector_body:\n    ld1h    { z1.h }, p0/z, [x0, x10, lsl #1]\n    ld1h    { z2.h }, p0/z, [x1, x10, lsl #1]\n    inch    x10\n    cmp     x9, x10\n    fdot    z0.s, z2.h, z1.h\n    b.ne    .Lvector_body\n\n    ptrue   p0.s\n    faddv   s0, p0, z0.s\n```\n\nSee the full [Compiler Explorer example](https://godbolt.org/z/MWsr9Er5v).\n\nPartial reductions now also work when the operation is predicated inside the loop. Code such as:\n\n``` js\nint64_t masked_dot(const int16_t *a, const int16_t *b,\n                   const int16_t *mask, int n) {\n  int64_t sum = 0;\n  for (int i = 0; i < n; ++i)\n    if (mask[i] > 0)\n      sum += (int64_t)a[i] * b[i];\n  return sum;\n}\n```\n\ncan stay vectorized while ensuring inactive lanes do not contribute to the result. LLVM creates predicates from the mask, uses them for the input loads, and performs the partial reduction with `sdot`:\n\n```\n    ld1h    { z2.h }, p0/z, [x2, x10, lsl #1]\n    cmpgt   p1.h, p0/z, z2.h, #0\n    ld1h    { z2.h }, p1/z, [x8, x10, lsl #1]\n    ld1h    { z3.h }, p1/z, [x1, x10, lsl #1]\n    inch    x10\n    sel     z2.h, p1, z2.h, z0.h\n    sdot    z1.d, z3.h, z2.h\n```\n\nThe [before-and-after output](https://godbolt.org/z/v5sf6YTMq) shows the full generated function.\n\nThe matching is no longer limited to a single, canonical `sum += product` form. LLVM 23 can handle subtraction performed in the middle block:\n\n```\nfor (int i = 0; i < n; ++i)\n  sum -= (int64_t)a[i] * b[i];\n```\n\nThe vector loop accumulates the products with `sdot`; the subtraction from the initial value is then performed in the middle block:\n\n```\n.Lvector_body:\n    ld1h    { z1.h }, p0/z, [x0, x10, lsl #1]\n    ld1h    { z2.h }, p0/z, [x1, x10, lsl #1]\n    inch    x10\n    sdot    z0.d, z2.h, z1.h\n    cmp     x9, x10\n    b.ne    .Lvector_body\n\n    ptrue   p0.d\n    uaddv   d0, p0, z0.d\n    fmov    x10, d0\n    sub     x2, x2, x10\n```\n\nand mixed add/subtract chains:\n\n```\nfor (int i = 0; i < n; ++i) {\n  sum += (int64_t)a[i] * b[i];\n  sum -= (int64_t)c[i] * d[i];\n}\n```\n\nIn the generated SVE loop, both products use `sdot`. The `subr` operations allow the additions and subtractions to be represented in the same partial-reduction chain:\n\n```\n    ld1h    { z1.h }, p0/z, [x0, x9, lsl #1]\n    ld1h    { z2.h }, p0/z, [x1, x9, lsl #1]\n    sdot    z0.d, z2.h, z1.h\n    ld1h    { z1.h }, p0/z, [x2, x9, lsl #1]\n    ld1h    { z2.h }, p0/z, [x3, x9, lsl #1]\n    subr    z0.d, z0.d, #0\n    sdot    z0.d, z2.h, z1.h\n    subr    z0.d, z0.d, #0\n```\n\nExplore the generated code for the [subtract reduction](https://godbolt.org/z/hczYh14hj) and the [add/sub chain](https://godbolt.org/z/fabqqrrq6).\n\nSupport has also been added for absolute-difference reductions, a common building block in image, video, and signal-processing code:\n\n``` js\nunsigned absdiff(const unsigned char *a, const unsigned char *b, int n) {\n  unsigned sum = 0;\n  for (int i = 0; i < n; ++i)\n    sum += __builtin_abs((int)a[i] - (int)b[i]);\n  return sum;\n}\n```\n\nLLVM can now recognize the widening absolute difference as part of a partial reduction instead of materializing a less efficient sequence. In this example, `uabd` calculates the byte-wise absolute differences and `udot` with a vector of ones accumulates them into 32-bit lanes:\n\n```\n.Lvector_body:\n    ldr     z2, [x12]\n    ldr     z3, [x13]\n    incb    x13\n    incb    x12\n    subs    x14, x14, x8\n    uabd    z2.b, p0/m, z2.b, z3.b\n    udot    z0.s, z2.b, z1.b\n    b.ne    .Lvector_body\n\n    ptrue   p0.s\n    uaddv   d0, p0, z0.s\n    fmov    w8, s0\n```\n\nSee the full [Compiler Explorer example](https://godbolt.org/z/9dcnodn7r).\n\n### Vectorizing find-last reductions\n\nBy Sander De Smalen\n\nConditional scalar assignments inside a loop often encode a “find last” operation:\n\n``` js\nint last_match(const int *cond, const int *values, int n, int needle) {\n  int result = -1;\n  for (int i = 0; i < n; ++i)\n    if (cond[i] == needle)\n      result = values[i];\n  return result;\n}\n```\n\nThese loops are reductions, even though they do not look like a sum or a minimum. LLVM 23 has better support for recognizing and vectorizing these conditional assignments as find-last reductions. The vector loop compares several elements at once and conditionally retains their indices:\n\n```\n.Lvector_body:\n    ld1w    { z2.s }, p1/z, [x0, x8, lsl #2]    // load cond[i]\n    cmpeq   p2.s, p1/z, z2.s, z1.s              // cond[i] == needle\n    csetm   x11, ne\n    whilelo p3.s, xzr, x11\n    ld1w    { z2.s }, p2/z, [x1, x8, lsl #2]    // load values[i]\n    incw    x8\n    mov     p0.b, p3/m, p2.b                    // if any cond[i] is true, then copy cond[] to p0\n    mov     z0.s, p3/m, z2.s                    // if any cond[i] is true, then copy values[] to z0\n    cmp     x10, x8\n    b.ne    .Lvector_body\n\n    mov     w8, #-1\n    cmp     x10, x9\n    clastb  w8, p0, w8, z0.s                    // extract last active lane, from values[]\n```\n\nThis pattern also appears in parsers, scans, and loops that retain the most recent qualifying value.\n\nSee the full [Compiler Explorer example](https://godbolt.org/z/TforMxb9d).\n\n### Explicit floating-point reduction builtins\n\nBy Sander De Smalen\n\nClang now provides two floating-point reduction builtins:\n\n```\nfloat ordered = __builtin_reduce_in_order_fadd(v, start);\nfloat fast    = __builtin_reduce_assoc_fadd(v, start);\n```\n\nThe optional second argument supplies the starting value; without it, the reduction starts from negative zero. The names make the key semantic choice explicit:\n\n- `__builtin_reduce_in_order_fadd` evaluates the reduction in lane order.\n- `__builtin_reduce_assoc_fadd` allows reassociation, giving the optimizer freedom to use a tree reduction or target-specific instructions.\n\nThis allows programmers to specify the required reduction semantics without applying broad fast-math options to the rest of the translation unit.\n\n### Improved interleaving and SVE shuffles\n\nBy Sander De Smalen\n\nInterleaved data is common in pixels, complex values, structures of channels, and packed sensor input. A lower vectorization factor may be appropriate for a loop, but previously it could prevent LLVM from selecting an interleaved memory operation.\n\nLLVM 23 allows interleave shuffles at these lower vectorization factors for SVE, including cases where a `tbl` sequence efficiently combines deinterleaving with zero-extension. In suitable kernels, one table lookup per output replaces a longer sequence of unpack and shuffle instructions:\n\n```\nstruct Pixel { unsigned short blue, green, red, alpha; };\n\nvoid sum_channels(const Pixel *__restrict pixels,\n                  double *__restrict sums, unsigned long count) {\n  double blue = 0, green = 0, red = 0, alpha = 0;\n  for (unsigned long i = 0; i < count; ++i) {\n    blue  += pixels[i].blue;\n    green += pixels[i].green;\n    red   += pixels[i].red;\n    alpha += pixels[i].alpha;\n  }\n  sums[0] = blue; sums[1] = green;\n  sums[2] = red;  sums[3] = alpha;\n}\n```\n\nHere, converting 16-bit channels to `double` makes the natural SVE vectorization factor smaller than the four-channel interleave factor. With `-O3 -ffast-math -march=armv9-a+sve2`, LLVM vectorizes the loop at `vscale x 2` and folds both the deinterleave and zero-extension into the table lookups:\n\n```\ntbl     z8.h, { z28.h }, z24.h\ntbl     z9.h, { z29.h }, z24.h\nucvtf   z8.d, p0/m, z8.d\nucvtf   z9.d, p0/m, z9.d\n```\n\nSee the full [Compiler Explorer example](https://godbolt.org/z/e4Pob4aMc).\n\n### Improved multi-vector loads and stores\n\nBy Sander De Smalen\n\nSVE and SME provide multi-vector memory operations. LLVM 23 improves the use of the instruction's addressing modes. For example,\n\n``` js\n#include <arm_sve.h>\n\nvoid copy_vector_pair(unsigned char *dst, const unsigned char *src) {\n  svcount_t all = svptrue_c8();\n  svuint8x2_t data = svld1_vnum_u8_x2(all, src, -2);\n  svst1_vnum_u8_x2(all, dst, 14, data);\n}\n```\n\nLLVM 23 folds both offsets directly into the multi-vector instructions, avoiding separate pointer adjustments:\n\n```\nptrue   pn8.b\nld1b    { z0.b, z1.b }, pn8/z, [x1, #-2, mul vl]\nst1b    { z0.b, z1.b }, pn8,   [x0, #14, mul vl]\n```\n\nSee the full [Compiler Explorer example](https://godbolt.org/z/r6v6fYnGf)\n\nStreaming-mode for SME2 code also benefits from enabling sub-register liveness. For example, this kernel loads two rows as vector pairs, then transposes the tuple view for a dot product:\n\n``` js\n#include <arm_sme.h>\n\nvoid dot_two_rows(unsigned long stride, const unsigned char *ptr)\n    __arm_streaming __arm_inout(\"za\") {\n  svcount_t all = svptrue_c8();\n  svuint8x2_t row0 = svld1_u8_x2(all, ptr);\n  svuint8x2_t row1 = svld1_u8_x2(all, ptr + stride);\n  svuint8x2_t lhs = svcreate2_u8(svget2_u8(row0, 0), svget2_u8(row1, 0));\n  svuint8x2_t rhs = svcreate2_u8(svget2_u8(row0, 1), svget2_u8(row1, 1));\n  svdot_za32_u8_vg1x2(0, lhs, rhs);\n}\n```\n\nLLVM 22 selects strided multi-vector loads, but must copy half of each tuple into consecutive registers for `udot`:\n\n```\nld1b    { z16.b, z24.b }, pn8/z, [x1]\nld1b    { z17.b, z25.b }, pn8/z, [x1, x0]\nmov     z0.d, z24.d\nmov     z1.d, z25.d\nudot    za.s[w8, 0, vgx2], { z16.b, z17.b }, { z0.b, z1.b }\n```\n\nLLVM 23 can keep both tuple views live and removes the copies:\n\n```\nld1b    { z16.b, z24.b }, pn8/z, [x1]\nld1b    { z17.b, z25.b }, pn8/z, [x1, x0]\nudot    za.s[w8, 0, vgx2], { z16.b, z17.b }, { z24.b, z25.b }\n```\n\nSee the full [Compiler Explorer example](https://godbolt.org/z/aacj3jocW)\n\n### Improved use of SVE immediates\n\nBy Sander De Smalen\n\nLLVM 23 recognizes more opportunities to use SVE instructions that have an immediate form, instead of first materializing a constant in a vector register and using an equivalent NEON instruction that lacks an immediate form. That saves an instruction and, just as importantly, one temporary register.\n\nFor example:\n\n```\n#include <arm_neon.h>\n\nuint32x4_t increment(uint32x4_t x) {\n  return vaddq_u32(x, vdupq_n_u32(1));\n}\n```\n\nLLVM 22 materialized the splat in a second vector register before using a NEON add:\n\n```\nmovi    v1.4s, #1\nadd     v0.4s, v0.4s, v1.4s\n```\n\nLLVM 23 can instead use the SVE immediate form directly:\n\n```\nadd     z0.s, z0.s, #1\n```\n\nSee the [Compiler Explorer example](https://godbolt.org/z/x5a5138vv)\n\n### Additional instruction-selection improvements\n\nBy Sander De Smalen\n\nLLVM 23 also includes local AArch64 combines that remove redundant instructions.\n\nOne example folds a condition-code materialization followed by a bit-test branch:\n\n```\ncset    w8, eq\ntbnz    w8, #0, .Lmatch\n```\n\ndirectly into the equivalent conditional branch:\n\n```\nb.eq    .Lmatch\n```\n\nAlthough individually small, these folds can reduce code size and leave fewer instructions for later scheduling stages.\n\n### Improved carry-less multiplication\n\nBy Matthew Devereau\n\nLLVM 23 improves AArch64 code generation for carry-less multiplication, an operation commonly used in cryptography and data-integrity algorithms. More fixed-width and scalable-vector forms are now mapped to sequences that utilize `PMUL`/` PMULL` instructions when SVE2/SME and AES extensions are available and scalar forms can now utilize the vector register. The cost model has also been updated so that LLVM can make better optimization decisions. \n\n### Faster compilation with GlobalISel\n\nBy Cullen Rhodes\n\nCompile-time has improved significantly in LLVM 23. Debug builds (`-O0 -g`) are 15% faster on CTMark, with some workloads such as SQLite almost 25% faster. Optimized builds (`-O3`) are also 8.7% faster. This is largely due to extensive improvements in `GlobalISel`, the default instruction selector at `-O0`, as well as broader improvements in Clang debug info and LLVM code generation.\n\n## Tools improvements\n\n### BOLT improvements\n\nBy Paschalis Mpeis\n\nLLVM 23 improves BOLT support for AArch64 binaries, with contributions focused on Branch Target Identification (BTI), compare-and-branch instructions, long-jump code layout, and usability.\n\nWe introduced initial BOLT support for processing BTI-enabled AArch64 binaries. BOLT can now check and patch PLT entries and indirect branch targets where possible. It can disassemble PLT entries in BTI binaries, patch LLD-generated PLTs with BTI landing pads, and patch ignored functions in place when indirect branches target them.\n\nWe also added compact-code-model support for Armv9.6-A `FEAT_CMPBR` compare-and-branch instructions. This includes support for block reordering, function splitting, branch inversion where legal, and trampoline handling for difficult branch cases.\n\nThis release also improves AArch64 usability. Unsupported or target-specific BOLT passes now report clear errors instead of crashing or silently doing nothing, and new documentation records AArch64 optimization flag support. BOLT also gains support for `SHT_CREL` code relocations, reports compressed debug sections cleanly, improves `--hugify` long-jump layout handling, and expands Arm test coverage.\n\n### Memory Protection Keys and Permission Overlays support in LLDB\n\nBy David Spickett\n\nMemory permissions are typically set in the page tables. Programs often change these permissions at runtime. For example they might initialise key structures and then make them read-only to protect them against attackers trying to modify them later.\n\nHowever, changing page table permissions is expensive. On Linux it is usually done with the `mprotect` system call. Linux has an alternative called [\"Memory Protection Keys\"](https://docs.kernel.org/core-api/protection-keys.html) which allows permissions to be changed without a system call. These keys work together with a hardware feature which stores \"permission overlays\". On AArch64 this is the Permission Overlay Extension (`FEAT_S1POE)` introduced in Armv8.8.\n\nNew in LLDB 23 is support for debugging Memory Protection Keys and permission overlays on AArch64 Linux.\n\nBelow is an example of a protection key fault:\n\n```\n(lldb) c\nProcess 462 resuming\nProcess 462 stopped\n<...> stop reason = signal SIGSEGV: failed protection key checks (fault address=0xffffff7d60000)\n<...>\n-> 106       read_only_page[0] = '?';\n```\n\nThis program tried to write to memory and was not allowed to do so. You can tell that the page table entry did have write permissions because of the description of the `SIGSEGV` signal. It refers specifically to protection keys (`SEGV_PKUERR)` , as opposed to \"invalid permissions for mapped object\" (`SEGV_ACCERR)` which is generated when the page table permissions are the cause.\n\nYou can check the permissions with the `memory region` command:\n\n```\n(lldb) memory region read_only_page\n[0x000ffffff7d60000-0x000ffffff7d70000) rw-\nprotection key: 6 (r--, effective: r--)\n```\n\nIn prior versions of LLDB, all you would see is the `rw-` on the 2nd line. This represents the permissions in the page table. They allow reading (`r`) and writing (` w`), but not execution (` x`).\n\nNew in LLDB 23 is the 3rd line. It shows:\n\n- The protection key assigned to this memory mapping (`6` ).\n- The permission overlay that the key refers to (`r--)` .\n- The effective permissions (`r--)` after combining the page table permissions with the permission overlay.\n\nPermission overlays can only remove permissions. In this example `rw-` was overlaid with `r--` . This overlay keeps the `r` permission but removes the `w` and `x` permissions. That is why the program failed to write. This memory mapping is read-only.\n\nLLDB also lets you access the `POR_EL0` register where the permission overlays are stored. To read about this and other AArch64 Linux features, see the [LLDB documentation](https://lldb.llvm.org/use/aarch64-linux.html).\n\n### Flang improvements\n\nBy Tom Eccles\n\nArm's contributions to LLVM 23 improved Flang's compatibility, performance and OpenMP support. Taskloop lowering is now more robust, with better handling of loop bounds, steps and privatized character values. Flang also supports the non-standard RTC intrinsic used by some existing Fortran applications. A new opt-in mode (`-freal-sum-association` in LLVM 23 or `-ffp-sum-association` in LLVM 24) can reassociate real-valued sums where permitted by the Fortran standard, exposing additional optimization opportunities while preserving explicit parentheses. Our biggest contribution was continuing upstream maintenance for flang's internal MLIR dialects, OpenMP on CPU support, and Windows support. I was the most active reviewer for flang OpenMP support (approximately 18% of reviews), and ranked third in flang as a whole (approximately 9% of reviews).\n\n### BOLT optimization for Flang\n\nBy Pawel Osmialowski\n\nThe purpose of BOLTing (optimizing with BOLT) the compiler itself is to make it compile the code faster; this is not about the optimization capabilities of the compiler itself (a common misconception).\n\nFollowing the already existing methodology for having the Clang compiler BOLTed, we have contributed a similar method for BOLTing the Flang compiler. As with Clang, our approach is also based on the CMake caches and reuses the in-tree profiling utilities that were already used for BOLTing Clang. The proposed CMake caches introduce two optimization methods: BOLT and BOLT+PGO (with and without LTO). We did our performance evaluation of the BOLT+PGO method on different servers using three Fortran codebase examples (the DBCSR library, the Polyhedron benchmark, and the LAPACK library) and saw the following speedup in compilation times:\n\n- DBCSR: 12.9 – 21.9%\n- Polyhedron (pb11): 12.7 – 18.7%\n- LAPACK-3.6.1: 15.3 – 16.2%\n\nDuring this work, however, we also discovered that the PGO-based optimization process is affected by a race condition that can cause the performance-training stage to hang intermittently. Until this issue is resolved, we recommend using BOLT without PGO when optimizing Flang.\n\n### Windows on Arm and Arm64EC support\n\nBy David Truby\n\nLLVM 23 extends Windows on Arm support across OpenMP, Flang and native interoperability. The OpenMP runtime now supports Arm64EC and Arm64X configurations, making it easier to combine Arm-native and emulated code. Flang’s runtime gains improved Windows implementations for date and time queries and file handling, together with Arm64EC build fixes. LLVM can also generate Arm64EC interoperability thunks for functions using `bfloat16` values.\n\n## Libraries improvements\n\n### Optimized AArch32 software floating point functions\n\nBy Simon Tatham\n\nIn the `compiler-rt` builtins library, there is now a full set of highly optimized implementations of floating-point basic arithmetic, for AArch32 systems supporting either of the Arm or Thumb2 instruction set but with no hardware FPU. A small number of functions also have optimized Thumb1 implementations, suitable for Armv6-M or Armv8-M Baseline. On average, the new functions run roughly twice as fast as the generic functions they replaced.\n\n### Libc vector math\n\nBy Dylan Fleming\n\nLLVM 23 introduces a `mathvec` component into the LLVM libc, providing LLVM vector math implementations of standard math functions that can ship alongside their scalar counterparts. Current work has focused on establishing the component structure, testing infrastructure, and delivering proofs of concept for both a cross-architecture (generic) and an AdvSIMD-optimised single-precision exponential (`expf`).\n\nLike scalar routines, vector routines are correctly rounded: they return the floating-point value nearest to the exact mathematical result for the supported rounding mode (presently, round to nearest, with ties to even). While this level of accuracy may be excessive for many use cases, correct rounding offers additional benefits.\n\nCorrectly rounded functions preserve properties of mathematical functions that lower-accuracy implementations may lose, potentially leading to unintuitive outputs, logical errors, and poor numerical stability. Among the properties frequently required by applications, monotonicity is preserved. For example, `x <= y` implies `f(x) <= f(y)` holds for floating-point quantities too. Perhaps more noticeably, functions remain within their true bounds, such as `|sin(x)| <= 1` . Correct rounding also guarantees bitwise-identical results across:\n\n- Scalar and vector implementations, enabling safe auto-vectorization without changing numerical behavior.\n- Architectures and platforms, improving reproducibility and avoiding target-dependent production results.\n- Library versions, helping users adopt the latest version and optimizations seamlessly.\n\nGoing forward, we plan to increase the coverage of vector routines, initially focusing on FP32 and FP16 before expanding to FP64 and BF16. These will first be delivered in generic form and enabled on all supported architectures, while target-specific optimizations will follow where observable gains can be achieved.\n\n## MLIR Improvements\n\n### TOSA\n\nBy Luke Hutton\n\nThe TOSA MLIR dialect in LLVM 23 continues to support TOSA 1.0 while introducing early support for new features from the TOSA 1.1 draft [specification](https://github.com/arm/tosa-specification).\n\nSupport for new specification features includes:\n\n- Introduction of block scaled tensor types\n\nBlock-scaled tensor types represent tensors whose values are grouped into blocks, with each block sharing a scale factor. The dialect supports the [OCP microscaling](https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf) (MX) formats, currently using 32-element blocks along the innermost dimension. The new type is supported by operations including `tosa.cast` and `tosa.const`; support across further operations aligned with the `EXT-MX-*` extensions is ongoing.\n\n- Dynamic shape expression and inference support\n\nExperimental `EXT-SHAPE` support allows TOSA graphs with unknown dimension sizes to be expressed. Once input shapes are known, shape inference can fold these expressions and propagate static shapes through the graph. In the current specification draft, shape values must resolve to constants during backend compilation; runtime-dynamic shapes are not yet supported.\n\n- Downgrade specification version\n\nA new best-effort transformation rewrites constructs available only in TOSA 1.1.draft into TOSA 1.0-compatible forms where possible. Because not every construct can be downgraded, the resulting IR should subsequently be validated against TOSA 1.0. This helps decouple front-end producers from back-end consumers supporting different TOSA versions.\n\n``` php\nfunc.func @main(%arg0: tensor<*xi1>) -> tensor<*xf32> {\n  %0 = tosa.cast %arg0 : (tensor<*xi1>) -> tensor<*xf32>\n  return %0 : tensor<*xf32>\n}\n\n$ mlir-opt --tosa-downgrade-1-1-to-1-0 test.mlir\n\nfunc.func @main(%arg0: tensor<*xi1>) -> tensor<*xf32> {\n  %0 = tosa.cast %arg0 : (tensor<*xi1>) -> tensor<*xi8>\n  %1 = tosa.cast %0 : (tensor<*xi8>) -> tensor<*xf32>\n  return %1 : tensor<*xf32>\n}\n```\n\n- TOSA to SPIR-V TOSA lowering\n\nA new lowering path that converts TOSA IR to SPIR-V dialect has been added. The lowering produces SPIR-V ARM Graph and Tensor extensions together with the SPIR-V TOSA extended instruction set. Unlike paths through lower-level dialects such as Linalg, Tensor, and Vector, it preserves higher-level TOSA semantics for backends that natively consume SPIR-V Graph and TOSA extensions.\n\n### SPIR-V\n\nBy Davide Grohmann\n\nLLVM 23 substantially expands machine learning support in MLIR's SPIR-V dialect. The new functionality provides tensor and graph representations, a comprehensive set of TOSA operations, low-precision floating-point types, and an end-to-end lowering path from the TOSA dialect.\n\n- Tensor and graph support\n\nLLVM 23 adds support for the [SPV_ARM_tensors](https://github.khronos.org/SPIRV-Registry/extensions/ARM/SPV_ARM_tensors.html) and [SPV_ARM_graph](https://github.khronos.org/SPIRV-Registry/extensions/ARM/SPV_ARM_graph.html) extensions.\n\nSPV_ARM_tensors introduces SPIR-V tensor types, including ranked and unranked tensors. SPV_ARM_graph represents dataflow computations over these tensors, including graph entry points, inputs, outputs, and constants. The SPIR-V dialect supports parsing, verification, and binary serialization and deserialization for both extensions.\n\nFor example, a graph operating on tensors can be represented as:\n\n- TOSA Extended Instruction Set\n\nLLVM 23 adds support for version 001000.1 of the [SPIR-V TOSA Extended Instruction Set](https://github.khronos.org/SPIRV-Registry/extended/TOSA.001000.1.html).\n\nIt covers operations for convolution, pooling, matrix multiplication, activation functions, elementwise arithmetic, reductions, data layout, gather and scatter, resize, casting, and rescaling.\n\nThese operations preserve TOSA semantics in SPIR-V instead of decomposing them into lower-level arithmetic and control-flow operations. This gives backends that understand TOSA operations more information when compiling an ML workload.\n\n- TOSA-to-SPIR-V lowering\n\nA new lowering path converts the MLIR TOSA dialect directly to the SPIR-V dialect. It combines SPV_ARM_tensors, SPV_ARM_graph, and the TOSA Extended Instruction Set to retain the graph structure, tensor types, and high-level operations of the original program.\n\nFor example, the following TOSA function:\n\ncan be lowered with: `mlir-opt --tosa-to-spirv-tosa input.mlir`\n\nThe resulting SPIR-V dialect retains both the graph and the ArgMax operation:\n\nThe lowering also handles graph constants and automatically derives the SPIR-V extensions and capabilities required by the generated module.\n\n- Float8 support\n\nLLVM 23 adds [SPV_EXT_float8](https://github.khronos.org/SPIRV-Registry/extensions/EXT/SPV_EXT_float8.html) support to the SPIR-V dialect, including the Float8E4M3EXT and Float8E5M2EXT formats. Float8 values can be used in scalar, vector, cooperative-matrix, and Arm tensor types.\n\nThe TOSA-to-SPIR-V path can therefore preserve Float8 tensor element types, enabling compact data representation for ML workloads without first promoting values to a wider floating-point type.\n\n- Replicated composite constants\n\nLLVM 23 adds support for [SPV_EXT_replicated_composites](https://github.khronos.org/SPIRV-Registry/extensions/EXT/SPV_EXT_replicated_composites.html). This extension represents composite constants whose elements all have the same value without listing every element individually.\n\nThe TOSA-to-SPIR-V lowering uses this representation for splat tensor constants, producing more compact SPIR-V modules.\n\n- Graph debugging information\n\nLLVM 23 also adds support for the [NonSemantic.Graph.DebugInfo.1](https://github.khronos.org/SPIRV-Registry/nonsemantic/NonSemantic.Graph.DebugInfo.html) extended instruction set. It represents debugging information for graphs, operations, and tensors without affecting program semantics.\n\nThe information is preserved during SPIR-V serialization and deserialization, helping tools relate operations in a generated SPIR-V graph to their compiler representation.\n\n- Experimental ML operations\n\nLLVM 23 supports the [Arm.ExperimentalMLOperations](https://github.khronos.org/SPIRV-Registry/extended/Arm.ExperimentalMLOperations.html).1 extended instruction set. It provides a generic representation for experimental or target-specific ML operations that are not part of the standard TOSA instruction set.\n\nThe TOSA-to-SPIR-V lowering can map selected TOSA custom operations to these instructions. This provides an extension point for introducing new operations while their interfaces and semantics are still evolving.\n\nRe-use is only permitted for informational and non-commercial or personal use only.", "url": "https://wpnews.pro/news/what-is-new-in-llvm-23", "canonical_source": "https://developer.arm.com/community/arm-community-blogs/b/tools-software-ides-blog/posts/what-is-new-in-llvm-23", "published_at": "2026-10-03 20:03:30+00:00", "updated_at": "2026-10-03 20:36:18.393023+00:00", "lang": "en", "topics": ["developer-tools", "ai-chips", "ai-infrastructure"], "entities": ["LLVM 23.1.0", "Arm", "Arm AGI CPU", "Armv9.2-A", "Armv9.6-A", "Armv9.7-A", "Clang", "MLIR"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-is-new-in-llvm-23", "markdown": "https://wpnews.pro/news/what-is-new-in-llvm-23.md", "text": "https://wpnews.pro/news/what-is-new-in-llvm-23.txt", "jsonld": "https://wpnews.pro/news/what-is-new-in-llvm-23.jsonld"}}