PyCuTe: Reference implementation and examples of the CuTe Layout NVIDIA researcher Cris Cecka released PyCuTe, a pure-Python reference implementation of the CuTe layout algebra used in CUTLASS 3.x and the CuTe DSL, enabling learning, prototyping, and test-vector generation without a GPU. The library implements all operations from the CuTe Whitepaper, including coalesce, composition, complement, and logical_divide, with Python 3.10+ as the only hard requirement and no third-party dependencies for the core algebra. A pure-Python reference implementation of CuTe — the hierarchical layout-and-tensor algebra at the heart of CUTLASS 3.x and the CuTe DSL. No GPU required. Where C++ CuTe is a header-only template library tightly coupled to CUDA, PyCuTe is plain Python you can import from any script — making it the place to learn the algebra, prototype new transformations, and generate test vectors for the C++ and DSL implementations. It implements the layout algebra from the CuTe Whitepaper — coalesce , composition , complement , logical divide , logical product , right inverse , left inverse , nullspace , recast , layout add , and greatest common domain — over integer and coordinate ArithTuple /basis strides, plus limited support for F2 XOR-swizzle strides. A thin Tensor / Accessor layer provides a reference data model. Cris Cecka, CuTe Layout Representation and Algebra, arXiv:2603.02298 . The Whitepaper is the authoritative source for every definition and post-condition; PyCuTe defers to it throughout. PyCuTe is installed in-place from a source checkout. Because many systems mark the system Python as externally managed PEP 668 , the recommended path is a virtual environment: python3 -m venv .venv source .venv/bin/activate pip install -e . core layout algebra only no third-party deps pip install -e ". viz " + visualization helpers svgwrite, tabulate pip install -e ". test " + everything needed to run the test suite Python 3.10+ is the only hard requirement, and the core algebra has no third-party dependencies. The optional extras add svgwrite / tabulate viz , sympy symbolic , and pytest test ; the draw latex helpers need only a LaTeX install e.g. TeX Live's pdflatex for their PDF step. You can also use the package straight from a repository checkout without installing — import pycute works as long as the repo is on PYTHONPATH . from pycute import A = Layout 3, 4 , 4, 1 3x4 row-major matrix A 2, 3 call the layout on a coordinate 11 A 11 a 1-D coordinate works too 11 size A , rank A 12, 2 coalesce Layout 2, 1, 6 , 1, 6, 2 Layout 12, 1 composition Layout 12 , Layout 4, 3 Layout 4, 3 , 1, 4 logical divide Layout 24 , Layout 4, 2 Layout 4, 2, 3 , 2, 1, 8 Build a Tensor and read/write data: T = make tensor Layout 4, 4 , 4, 1 4x4 row-major T 1, 2 = 42.0 T 1, 2 42.0 Print or draw a layout see Visualization visualization : python from pycute.util import print tensor, draw svg print tensor Layout 4, 8 , 1, 4 4, 8 : 1, 4 0 4 8 12 16 20 24 28 1 5 9 13 17 21 25 29 2 6 10 14 18 22 26 30 3 7 11 15 19 23 27 31 draw svg Layout 4, 8 , 1, 4 Saved as layout.svg The whole of CuTe fits in one sentence: A is a function from coordinates to offsets, defined by a Layout the coordinate s domain and a Shape how coordinates become offsets of the same hierarchical profile. Stride The stride is what turns a shape into a row-major, column-major, or arbitrarily nested map: | Layout | Description | |---|---| Layout 4, 8 , 8, 1 | 4×8 row-major | Layout 4, 8 , 1, 4 | 4×8 column-major | Layout 2, 4 , 8 , 1, 16 , 2 | hierarchical nested modes | A small algebra combines layouts to express tiling, partitioning, vectorization, and layout analysis. Every operation is a pure function that takes layouts and returns another Layout : — simplify to the fewest modes with the same map. coalesce A — functional composition; index composition A, B A through B .— the "missing" modes that fill out complement A A 's codomain.— factor logical divide A, T A into tiles of shape T tiling .— replicate logical product A, B A 's pattern across B repetition .— invert and analyze maps. right inverse / left inverse / nullspace A Tensor is a Layout paired with an e.g. a pointer : evaluating it at a coordinate evaluates the layout to an offset and dereferences the accessor at that offset. That short story is the whole of CuTe — everything else is a refinement or application of it. For the careful treatment of each piece, read the Accessor documentation documentation . Start with docs/index.md /NVlabs/CuTe/blob/main/docs/index.md . The documentation builds up from hierarchical tuples to layouts to the full algebra, with runnable examples drawn from the unit tests: | File | Topic | |---|---| docs/00 quickstart.md | docs/01 htuple.md docs/02 shape stride.md Shape , Stride , and the integer-modules strides live in docs/03 layout.md Layout : construction, evaluation, coordinates, slicing docs/04 layout algebra.md docs/05 tensor.md Tensor and Accessor docs/06 swizzle.md Swizzle and F2 -stride layouts docs/07 visualization.md print tensor , draw svg , draw latex , and color functors docs/08 api reference.md PyCuTe renders layouts as ASCII tables print tensor , print table , colored SVGs draw svg , draw svg tv , or TikZ/PDF draw latex , draw latex tv — the analogue of cute::print latex . Every figure below is a plain PyCuTe Layout drawn with pycute.util ; regenerate them all with examples/readme figures.py /NVlabs/CuTe/blob/main/examples/readme figures.py : python -m examples.readme figures writes docs/images/ .svg needs the viz extra A layout is a shape plus a stride. The same 8×8 shape with two different strides gives a row-major or a column-major map; each cell is labeled with its offset and colored by offset % 8 : Layout 8,8 , 8,1 Row-major | Layout 8,8 , 1,8 Column-major | Layout 4,2 , 2,4 , 1,32 , 4,8 Blocked | Thread-value layouts. draw svg tv shows how a warp's thread, value pairs tile a matrix — here the C-accumulator of an SM80 16×8 MMA, colored by thread. This is precisely the partitioning the layout algebra produces: Layout 4,8 , 2,2 , 32,1 , 16,8 tid, vid → m, n in a 16×8 tile Swizzles are just F2 strides. Coloring a shared-memory tile by bank bank color 8x makes conflicts visible. A row-major tile stores every column in a single bank vertical stripes → 8-way conflict ; swapping the integer strides for F2 XOR strides permutes each row so that every column spans all 8 banks — conflict-free — with no special-casing in the algebra: Layout 8,8 , F2 1 ,F2 9 Swizzled Column-major | Layout 8,8 , F2 9 ,F2 1 Swizzled Row-major | See docs/07 visualization.md /NVlabs/CuTe/blob/main/docs/07 visualization.md for every drawer, the r, g, b color-functor catalog index grey 8x , bank color 32x , thread color 8x , … , and the LaTeX/PDF output. pycute/ ├── docs/ documentation start at docs/index.md ; figures in docs/images/ ├── examples/ standalone scripts einsum, TV-layout, README figures ├── test/ pytest unit tests one test .py per operation └── pycute/ the importable package └── util/ optional printing and visualization helpers The suite uses pytest https://docs.pytest.org . Install the test extra see Installation installation and run it from the repository root: pytest the whole suite, quiet pytest --log-cli-level DEBUG with live logging pytest test/test coalesce.py a single module pytest -k coalesce tests matching a keyword - Cris Cecka. CuTe Layout Representation and Algebra. arXiv:2603.02298 https://arxiv.org/abs/2603.02298 - Jack Carlisle, Jay Shah, Reuben Stern, Paul VanKoughnett. Categorical Foundations for CuTe Layouts. arXiv:2601.05972 https://arxiv.org/abs/2601.05972 - Yang Shi, U. N. Niranjan, Animashree Anandkumar, Cris Cecka. Tensor Contractions with Extended BLAS Kernels on CPU and GPU. HiPC 2016, pp. 193–202 https://ieeexplore.ieee.org/document/7839684 - Bastian Hagedorn, Bin Fan, Hanfeng Chen, Cris Cecka, Michael Garland, Vinod Grover. Graphene: An IR for Optimized Tensor Computations on GPUs. ASPLOS 2023, pp. 302–313 https://dl.acm.org/doi/10.1145/3582016.3582018 - NVIDIA CUTLASS / CuTe C++ https://github.com/NVIDIA/cutlass — the original C++ implementation. - NVIDIA CuTe DSL https://docs.nvidia.com/cutlass/latest/media/docs/pythonDSL/overview.html — the Python DSL that JIT-compiles CuTe kernels. Copyright c 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. SPDX-License-Identifier: Apache-2.0