NVIDIA today provided more details about its latest custom Armv9.2 "Vera" CPU, which the company describes as a pinnacle of its CPU engineering. Designed to coexist within NVL rackscale systems or operate independently, NVIDIA has seen significant interest from its customers in deploying the "Vera" CPU across enterprises. The company optimized its design for data analysis, high-performance single-threaded tasks, agentic AI, and large-scale software operations, covering a wide range of use cases. The entire "Olympus" CPU core is divided into four major areas: front end, mid-core, execution engine, and cache subsystems. In the front end, NVIDIA has developed a branch predictor that works actively with the instruction fetching unit to keep the core supplied with useful instructions. Tasks such as agent runtimes, compilers, and large data processing workloads often involve control-flow changes, which can leave execution units in the CPU idle. The "Olympus" core uses a 10-wide decode engine to ensure a continuous flow of instructions to execution units, keeping them fully utilized. To prepare the data path for execution, NVIDIA created a neural branch predictor that reduces the execution of incorrect paths, supporting up to two taken branches per cycle. This is crucial for control-flow changes when software execution on large data processing suddenly shifts.
NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI