We've created the first vectorized Quicksort
Google researchers built the first portable vectorized Quicksort, using Highway's portable SIMD functions to run on six instruction sets across three architectures, and reported a 9-19x speedup over t…
Google researchers built the first portable vectorized Quicksort, using Highway's portable SIMD functions to run on six instruction sets across three architectures, and reported a 9-19x speedup over t…
A new arXiv paper (2609.16338v1) introduces BITCOS, a distribution-adaptive storage layout for ternary LLMs that costs 2 − z bits per weight given a zero density z, after measuring 29 ternary models a…
A developer built SetrixDB, an exact set engine in Go that performs presence checks and intersections over uint64 IDs using a from-scratch Minimal Perfect Hash Function (CHD v2) and AVX-512-accelerate…
FFmpeg 9.0, the latest feature release of the open-source multimedia library, is now available with more Vulkan acceleration, including APV video decoding and Apple ProRes RAW Vulkan acceleration, plu…
Intel engineers posted initial GCC compiler patches for the AI Compute Extensions (ACE), a cross-vendor x86 specification for optimizing AI and machine learning workloads. The ACE specification, devel…
OpenCV and AMD announced a collaboration to accelerate computer vision and Vision AI workloads on AMD hardware, with AMD becoming an OpenCV 5 Launch Partner and an OpenCV Gold Sponsor. The partnership…
A developer built a CPU-only, distributed LLM pipeline to extract structured data from 10,000 full-text research papers, using a 35B MoE model running on a cluster of older x86 servers with zero GPUs.…
Based on the article, Geekbench 6 is a consumer-focused benchmark suite distributed in binary form that heavily utilizes modern CPU instruction set extensions like AVX-512 and AMX, particularly in wor…