GPGPU-SIM – Cycle-Level Simulator for Nvidia GPUs with CUDA or OpenCL Workloads GPGPU-Sim 4.0, a cycle-level simulator for Nvidia GPUs running CUDA or OpenCL workloads, is now compatible with the Accel-Sim simulation framework and can run NVIDIA SASS traces generated by NVBit, according to the project's release documentation. The release replaces the GPUWattch energy model used in versions 3.2.0 through 4.1.0 with AccelWattch version 1.0, which is validated against an NVIDIA Volta QV100 GPU. The simulator has been tested with subsets of CUDA versions 4.2, 5.0, 5.5, 6.0, 7.5, 8.0, 9.0, 9.1, 10, 11, and 12, and includes the AerialVision performance visualization tool. Welcome to GPGPU-Sim, a cycle-level simulator modeling contemporary graphics processing units GPUs running GPU computing workloads written in CUDA or OpenCL. Also included in GPGPU-Sim is a performance visualization tool called AerialVision and a configurable and extensible power model called AccelWattch. GPGPU-Sim and AccelWattch have been rigorously validated with performance and power measurements of real hardware GPUs. This version of GPGPU-Sim has been tested with a subset of CUDA version 4.2, 5.0, 5.5, 6.0, 7.5, 8.0, 9.0, 9.1, 10, 11, and 12 Please see the copyright notice in the file COPYRIGHT distributed with this release in the same directory as this file. GPGPU-Sim 4.0 is compatible with Accel-Sim simulation framework. With the support of Accel-Sim, GPGPU-Sim 4.0 can run NVIDIA SASS traces trace-based simulation generated by NVIDIA's dynamic binary instrumentation tool NVBit . For more information about Accel-Sim, see https://accel-sim.github.io/ https://accel-sim.github.io/ If you use GPGPU-Sim 4.0 in your research, please cite: Mahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G Rogers. Accel-Sim: An Extensible Simulation Framework for Validated GPU Modeling. In proceedings of the 47th IEEE/ACM International Symposium on Computer Architecture ISCA , May 29 - June 3, 2020. If you use CuDNN or PyTorch support execution-driven simulation , checkpointing or our new debugging tool for functional simulation errors in GPGPU-Sim for your research, please cite: Jonathan Lew, Deval Shah, Suchita Pati, Shaylin Cattell, Mengchi Zhang, Amruth Sandhupatla, Christopher Ng, Negar Goli, Matthew D. Sinclair, Timothy G. Rogers, Tor M. Aamodt Analyzing Machine Learning Workloads Using a Detailed GPU Simulator, arXiv:1811.08933, https://arxiv.org/abs/1811.08933 https://arxiv.org/abs/1811.08933 If you use the Tensor Core model in GPGPU-Sim or GPGPU-Sim's CUTLASS Library for your research please cite: Md Aamir Raihan, Negar Goli, Tor Aamodt, Modeling Deep Learning Accelerator Enabled GPUs, arXiv:1811.08309, https://arxiv.org/abs/1811.08309 https://arxiv.org/abs/1811.08309 If you use the AccelWattch power model in your research, please cite: Vijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan, Amogh Manjunath, Timothy G. Rogers, Tor M. Aamodt, and Nikos Hardavellas. 2021. AccelWattch: A Power Modeling Framework for Modern GPUs. In MICRO54: 54th Annual IEEE/ACM International Symposium on Microarchitecture MICRO ’21 , October 18–22, 2021, Virtual Event, Greece. If you use the support for CUDA dynamic parallelism in your research, please cite: Jin Wang and Sudhakar Yalamanchili, Characterization and Analysis of Dynamic Parallelism in Unstructured GPU Applications, 2014 IEEE International Symposium on Workload Characterization IISWC , November 2014. If you use figures plotted using AerialVision in your publications, please cite: Aaron Ariel, Wilson W. L. Fung, Andrew Turner, Tor M. Aamodt, Visualizing Complex Dynamics in Many-Core Accelerator Architectures, In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software ISPASS , pp. 164-174, White Plains, NY, March 28-30, 2010. This file contains instructions on installing, building and running GPGPU-Sim. Detailed documentation on what GPGPU-Sim models, how to configure it, and a guide to the source code can be found here: http://gpgpu-sim.org/manual/ http://gpgpu-sim.org/manual/ . Instructions for building doxygen source code documentation are included below. Previous versions of GPGPU-Sim 3.2.0 to 4.1.0 included the GPUWattch Energy model http://gpgpu-sim.org/gpuwattch/ which has been replaced by AccelWattch version 1.0 in GPGPU-Sim version 4.2.0. AccelWattch supports modern GPUs and is validated against a NVIDIA Volta QV100 GPU. Detailed documentation on AccelWattch can be found here: AccelWattch Overview https://github.com/VijayKandiah/accel-sim-framework accelwattch-overview and AccelWattch MICRO'21 Artifact Manual https://github.com/VijayKandiah/accel-sim-framework/blob/release/AccelWattch.md . If you have questions, please sign up for the google groups page see gpgpu-sim.org , but note that use of this simulator does not imply any level of support. Questions answered on a best effort basis. To submit a bug report, go here: http://www.gpgpu-sim.org/bugs/ http://www.gpgpu-sim.org/bugs/ See Section 2 "INSTALLING, BUILDING and RUNNING GPGPU-Sim" below to get started. See file CHANGES for updates in this and earlier versions. GPGPU-Sim was created by Tor Aamodt's research group at the University of British Columbia. Many have directly contributed to development of GPGPU-Sim including: Tor Aamodt, Wilson W.L. Fung, Ali Bakhoda, George Yuan, Ivan Sham, Henry Wong, Henry Tran, Andrew Turner, Aaron Ariel, Inderpret Singh, Tim Rogers, Jimmy Kwa, Andrew Boktor, Ayub Gubran Tayler Hetherington and others. GPGPU-Sim models the features of a modern graphics processor that are relevant to non-graphics applications. The first version of GPGPU-Sim was used in a MICRO'07 paper and follow-on ACM TACO paper on dynamic warp formation. That version of GPGPU-Sim used the SimpleScalar PISA instruction set for functional simulation, and various configuration files indicating which loops should be spawned as kernels on the GPU, along with reconvergence points required for SIMT execution to provide a programming model simlar to CUDA/OpenCL. Creating benchmarks for the original GPGPU-Sim simulator was a very time consuming process and the validity of code generation for CPU run on a GPU was questioned by some. These issues motivated the development an interface for directly running CUDA applications to leverage the growing number of applications being developed to use CUDA. We subsequently added support for OpenCL and removed all SimpleScalar code. The interconnection network is simulated using the booksim simulator developed by Bill Dally's research group at Stanford. To produce output that matches the output from running the same CUDA program on the GPU, we have implemented several PTX instructions using the CUDA Math library part of the CUDA toolkit . Code to interface with the CUDA Math library is contained in cuda-math.h, which also includes several structures derived from vector types.h one of the CUDA header files . AccelWattch introduced in GPGPU-Sim 4.2.0 was developed by researchers at Northwestern University, Purdue University, and the University of British Columbia. Contributors to AccelWattch include Nikos Hardavellas's research group at Northwestern University: Vijay Kandiah; Tor Aamodt's research group at the University of British Columbia: Scott Peverelle; and Timothy Rogers's research group at Purdue University: Mahmoud Khairy, Junrui Pan, and Amogh Manjunath. AccelWattch leverages McPAT, which was developed by Sheng Li et al. at the University of Notre Dame, Hewlett-Packard Labs, Seoul National University, and the University of California, San Diego. The McPAT paper can be found at http://www.hpl.hp.com/research/mcpat/micro09.pdf http://www.hpl.hp.com/research/mcpat/micro09.pdf . Assuming all dependencies required by GPGPU-Sim are installed on your system, to build GPGPU-Sim all you need to do is add the following line to your ~/.bashrc file assuming the CUDA Toolkit was installed in /usr/local/cuda : export CUDA INSTALL PATH=/usr/local/cuda then type bash source setup environment make If the above fails, see "Step 1" and "Step 2" below. If the above worked, see "Step 3" below, which explains how to run a CUDA benchmark on GPGPU-Sim. GPGPU-Sim was developed on SUSE Linux this release was tested with SUSE version 11.3 and has been used on several other Linux platforms both 32-bit and 64-bit systems . In principle, GPGPU-Sim should work with any linux distribution as long as the following software dependencies are satisfied. Download and install the CUDA Toolkit. It is recommended to use version 3.1 for normal PTX simulation and version 4.0 for cuobjdump support and/or to use PTXPlus Harware instruction set support . Note that it is possible to have multiple versions of the CUDA toolkit installed on a single system -- just install them in different directories and set your CUDA INSTALL PATH environment variable to point to the version you want to use. Optional If you want to run OpenCL on the simulator, download and install NVIDIA's OpenCL driver from http://developer.nvidia.com/opencl http://developer.nvidia.com/opencl . Update your PATH and LD LIBRARY PATH as indicated by the NVIDIA install scripts. Note that you will need to use the lib64 directory if you are using a 64-bit machine. We have tested OpenCL on GPGPU-Sim using NVIDIA driver version 256.40 http://developer.download.nvidia.com/compute/cuda/3 1/drivers/devdriver 3.1 linux 64 256.40.run http://developer.download.nvidia.com/compute/cuda/3 1/drivers/devdriver 3.1 linux 64 256.40.run This version of GPGPU-Sim has been updated to support more recent versions of the NVIDIA drivers tested on version 295.20 . GPGPU-Sim dependencies: - gcc - g++ - make - makedepend - xutils - bison - flex - zlib - CUDA Toolkit GPGPU-Sim documentation dependencies: - doxygen - graphvi AerialVision dependencies: - python-pmw - python-ply - python-numpy - libpng12-dev - python-matplotlib We used gcc/g++ version 4.5.1, bison version 2.4.1, and flex version 2.5.35. If you are using Ubuntu, the following commands will install all required dependencies besides the CUDA Toolkit. GPGPU-Sim dependencies: sudo apt-get install build-essential xutils-dev bison zlib1g-dev flex libglu1-mesa-dev GPGPU-Sim documentation dependencies: sudo apt-get install doxygen graphviz AerialVision dependencies: sudo apt-get install python-pmw python-ply python-numpy libpng12-dev python-matplotlib CUDA SDK dependencies: sudo apt-get install libxi-dev libxmu-dev libglut3-dev If you are running applications which use NVIDIA libraries such as cuDNN and cuBLAS, install them too. Finally, ensure CUDA INSTALL PATH is set to the location where you installed the CUDA Toolkit e.g., /usr/local/cuda and that $CUDA INSTALL PATH/bin is in your PATH. You probably want to modify your .bashrc file to incude the following this assumes the CUDA Toolkit was installed in /usr/local/cuda : export CUDA INSTALL PATH=/usr/local/cuda export PATH=$CUDA INSTALL PATH/bin If running applications which use cuDNN or cuBLAS: export CUDNN PATH=