{"slug": "gpgpu-sim-cycle-level-simulator-for-nvidia-gpus-with-cuda-or-opencl-workloads", "title": "GPGPU-SIM – Cycle-Level Simulator for Nvidia GPUs with CUDA or OpenCL Workloads", "summary": "GPGPU-Sim 4.0, a cycle-level simulator for Nvidia GPUs running CUDA or OpenCL workloads, is now compatible with the Accel-Sim simulation framework and can run NVIDIA SASS traces generated by NVBit, according to the project's release documentation. The release replaces the GPUWattch energy model used in versions 3.2.0 through 4.1.0 with AccelWattch version 1.0, which is validated against an NVIDIA Volta QV100 GPU. The simulator has been tested with subsets of CUDA versions 4.2, 5.0, 5.5, 6.0, 7.5, 8.0, 9.0, 9.1, 10, 11, and 12, and includes the AerialVision performance visualization tool.", "body_md": "Welcome to GPGPU-Sim, a cycle-level simulator modeling contemporary graphics processing units (GPUs) running GPU computing workloads written in CUDA or OpenCL. Also included in GPGPU-Sim is a performance visualization tool called AerialVision and a configurable and extensible power model called AccelWattch. GPGPU-Sim and AccelWattch have been rigorously validated with performance and power measurements of real hardware GPUs.\n\nThis version of GPGPU-Sim has been tested with a subset of CUDA version 4.2, 5.0, 5.5, 6.0, 7.5, 8.0, 9.0, 9.1, 10, 11, and 12\n\nPlease see the copyright notice in the file COPYRIGHT distributed with this release in the same directory as this file.\n\nGPGPU-Sim 4.0 is compatible with Accel-Sim simulation framework. With the support\nof Accel-Sim, GPGPU-Sim 4.0 can run NVIDIA SASS traces (trace-based simulation)\ngenerated by NVIDIA's dynamic binary instrumentation tool (NVBit). For more information\nabout Accel-Sim, see [https://accel-sim.github.io/](https://accel-sim.github.io/)\n\nIf you use GPGPU-Sim 4.0 in your research, please cite:\n\nMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G Rogers. Accel-Sim: An Extensible Simulation Framework for Validated GPU Modeling. In proceedings of the 47th IEEE/ACM International Symposium on Computer Architecture (ISCA), May 29 - June 3, 2020.\n\nIf you use CuDNN or PyTorch support (execution-driven simulation), checkpointing or our new debugging tool for functional simulation errors in GPGPU-Sim for your research, please cite:\n\nJonathan Lew, Deval Shah, Suchita Pati, Shaylin Cattell, Mengchi Zhang, Amruth Sandhupatla,\nChristopher Ng, Negar Goli, Matthew D. Sinclair, Timothy G. Rogers, Tor M. Aamodt\nAnalyzing Machine Learning Workloads Using a Detailed GPU Simulator, arXiv:1811.08933,\n[https://arxiv.org/abs/1811.08933](https://arxiv.org/abs/1811.08933)\n\nIf you use the Tensor Core model in GPGPU-Sim or GPGPU-Sim's CUTLASS Library for your research please cite:\n\nMd Aamir Raihan, Negar Goli, Tor Aamodt,\nModeling Deep Learning Accelerator Enabled GPUs, arXiv:1811.08309,\n[https://arxiv.org/abs/1811.08309](https://arxiv.org/abs/1811.08309)\n\nIf you use the AccelWattch power model in your research, please cite:\n\nVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan, Amogh Manjunath, Timothy G. Rogers, Tor M. Aamodt, and Nikos Hardavellas. 2021. AccelWattch: A Power Modeling Framework for Modern GPUs. In MICRO54: 54th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’21), October 18–22, 2021, Virtual Event, Greece.\n\nIf you use the support for CUDA dynamic parallelism in your research, please cite:\n\nJin Wang and Sudhakar Yalamanchili, Characterization and Analysis of Dynamic Parallelism in Unstructured GPU Applications, 2014 IEEE International Symposium on Workload Characterization (IISWC), November 2014.\n\nIf you use figures plotted using AerialVision in your publications, please cite:\n\nAaron Ariel, Wilson W. L. Fung, Andrew Turner, Tor M. Aamodt, Visualizing Complex Dynamics in Many-Core Accelerator Architectures, In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), pp. 164-174, White Plains, NY, March 28-30, 2010.\n\nThis file contains instructions on installing, building and running GPGPU-Sim.\nDetailed documentation on what GPGPU-Sim models, how to configure it, and a\nguide to the source code can be found here: [http://gpgpu-sim.org/manual/](http://gpgpu-sim.org/manual/).\nInstructions for building doxygen source code documentation are included below.\n\nPrevious versions of GPGPU-Sim (3.2.0 to 4.1.0) included the [GPUWattch Energy model](http://gpgpu-sim.org/gpuwattch/) which has been replaced by AccelWattch version 1.0 in GPGPU-Sim version 4.2.0. AccelWattch supports modern GPUs and is validated against a NVIDIA Volta QV100 GPU. Detailed documentation on AccelWattch can be found here: [AccelWattch Overview](https://github.com/VijayKandiah/accel-sim-framework#accelwattch-overview) and [AccelWattch MICRO'21 Artifact Manual](https://github.com/VijayKandiah/accel-sim-framework/blob/release/AccelWattch.md).\n\nIf you have questions, please sign up for the google groups page (see gpgpu-sim.org), but note that use of this simulator does not imply any level of support. Questions answered on a best effort basis.\n\nTo submit a bug report, go here: [http://www.gpgpu-sim.org/bugs/](http://www.gpgpu-sim.org/bugs/)\n\nSee Section 2 \"INSTALLING, BUILDING and RUNNING GPGPU-Sim\" below to get started.\n\nSee file CHANGES for updates in this and earlier versions.\n\nGPGPU-Sim was created by Tor Aamodt's research group at the University of British Columbia. Many have directly contributed to development of GPGPU-Sim including: Tor Aamodt, Wilson W.L. Fung, Ali Bakhoda, George Yuan, Ivan Sham, Henry Wong, Henry Tran, Andrew Turner, Aaron Ariel, Inderpret Singh, Tim Rogers, Jimmy Kwa, Andrew Boktor, Ayub Gubran Tayler Hetherington and others.\n\nGPGPU-Sim models the features of a modern graphics processor that are relevant to non-graphics applications. The first version of GPGPU-Sim was used in a MICRO'07 paper and follow-on ACM TACO paper on dynamic warp formation. That version of GPGPU-Sim used the SimpleScalar PISA instruction set for functional simulation, and various configuration files indicating which loops should be spawned as kernels on the GPU, along with reconvergence points required for SIMT execution to provide a programming model simlar to CUDA/OpenCL. Creating benchmarks for the original GPGPU-Sim simulator was a very time consuming process and the validity of code generation for CPU run on a GPU was questioned by some. These issues motivated the development an interface for directly running CUDA applications to leverage the growing number of applications being developed to use CUDA. We subsequently added support for OpenCL and removed all SimpleScalar code.\n\nThe interconnection network is simulated using the booksim simulator developed by Bill Dally's research group at Stanford.\n\nTo produce output that matches the output from running the same CUDA program on the GPU, we have implemented several PTX instructions using the CUDA Math library (part of the CUDA toolkit). Code to interface with the CUDA Math library is contained in cuda-math.h, which also includes several structures derived from vector_types.h (one of the CUDA header files).\n\nAccelWattch (introduced in GPGPU-Sim 4.2.0) was developed by researchers at Northwestern University, Purdue University, and the University of British Columbia. Contributors to AccelWattch include Nikos Hardavellas's research group at Northwestern University: Vijay Kandiah; Tor Aamodt's research group at the University of British Columbia: Scott Peverelle; and Timothy Rogers's research group at Purdue University: Mahmoud Khairy, Junrui Pan, and Amogh Manjunath.\n\nAccelWattch leverages McPAT, which was developed by Sheng Li et al. at the\nUniversity of Notre Dame, Hewlett-Packard Labs, Seoul National University, and\nthe University of California, San Diego. The McPAT paper can be found at\n[http://www.hpl.hp.com/research/mcpat/micro09.pdf](http://www.hpl.hp.com/research/mcpat/micro09.pdf).\n\nAssuming all dependencies required by GPGPU-Sim are installed on your system, to build GPGPU-Sim all you need to do is add the following line to your ~/.bashrc file (assuming the CUDA Toolkit was installed in /usr/local/cuda):\n\n```\n  export CUDA_INSTALL_PATH=/usr/local/cuda\n```\n\nthen type\n\n```\n  bash\n  source setup_environment\n  make\n```\n\nIf the above fails, see \"Step 1\" and \"Step 2\" below.\n\nIf the above worked, see \"Step 3\" below, which explains how to run a CUDA benchmark on GPGPU-Sim.\n\nGPGPU-Sim was developed on SUSE Linux (this release was tested with SUSE version 11.3) and has been used on several other Linux platforms (both 32-bit and 64-bit systems). In principle, GPGPU-Sim should work with any linux distribution as long as the following software dependencies are satisfied.\n\nDownload and install the CUDA Toolkit. It is recommended to use version 3.1 for normal PTX simulation and version 4.0 for cuobjdump support and/or to use PTXPlus (Harware instruction set support). Note that it is possible to have multiple versions of the CUDA toolkit installed on a single system -- just install them in different directories and set your CUDA_INSTALL_PATH environment variable to point to the version you want to use.\n\n[Optional] If you want to run OpenCL on the simulator, download and install\nNVIDIA's OpenCL driver from [http://developer.nvidia.com/opencl](http://developer.nvidia.com/opencl). Update your\nPATH and LD_LIBRARY_PATH as indicated by the NVIDIA install scripts. Note that\nyou will need to use the lib64 directory if you are using a 64-bit machine. We\nhave tested OpenCL on GPGPU-Sim using NVIDIA driver version 256.40\n[http://developer.download.nvidia.com/compute/cuda/3_1/drivers/devdriver_3.1_linux_64_256.40.run](http://developer.download.nvidia.com/compute/cuda/3_1/drivers/devdriver_3.1_linux_64_256.40.run)\nThis version of GPGPU-Sim has been updated to support more recent versions of\nthe NVIDIA drivers (tested on version 295.20).\n\nGPGPU-Sim dependencies:\n\n- gcc\n- g++\n- make\n- makedepend\n- xutils\n- bison\n- flex\n- zlib\n- CUDA Toolkit\n\nGPGPU-Sim documentation dependencies:\n\n- doxygen\n- graphvi\n\nAerialVision dependencies:\n\n- python-pmw\n- python-ply\n- python-numpy\n- libpng12-dev\n- python-matplotlib\n\nWe used gcc/g++ version 4.5.1, bison version 2.4.1, and flex version 2.5.35.\n\nIf you are using Ubuntu, the following commands will install all required dependencies besides the CUDA Toolkit.\n\nGPGPU-Sim dependencies:\n\n```\nsudo apt-get install build-essential xutils-dev bison zlib1g-dev flex libglu1-mesa-dev\n```\n\nGPGPU-Sim documentation dependencies:\n\n```\nsudo apt-get install doxygen graphviz\n```\n\nAerialVision dependencies:\n\n```\nsudo apt-get install python-pmw python-ply python-numpy libpng12-dev python-matplotlib\n```\n\nCUDA SDK dependencies:\n\n```\nsudo apt-get install libxi-dev libxmu-dev libglut3-dev\n```\n\nIf you are running applications which use NVIDIA libraries such as cuDNN and cuBLAS, install them too.\n\nFinally, ensure CUDA_INSTALL_PATH is set to the location where you installed the CUDA Toolkit (e.g., /usr/local/cuda) and that $CUDA_INSTALL_PATH/bin is in your PATH. You probably want to modify your .bashrc file to incude the following (this assumes the CUDA Toolkit was installed in /usr/local/cuda):\n\n```\nexport CUDA_INSTALL_PATH=/usr/local/cuda\nexport PATH=$CUDA_INSTALL_PATH/bin\n```\n\nIf running applications which use cuDNN or cuBLAS:\n\n```\nexport CUDNN_PATH=<Path To cuDNN Directory>\nexport LD_LIBRARY_PATH=$CUDA_INSTALL_PATH/lib64:$CUDA_INSTALL_PATH/lib:$CUDNN_PATH/lib64\n```\n\nYou can also opt in the prebuilt images on [https://github.com/accel-sim/Dockerfile/pkgs/container/accel-sim-framework](https://github.com/accel-sim/Dockerfile/pkgs/container/accel-sim-framework).\n\n```\n# Pull Ubuntu 24.04 with Cuda 12.8\ndocker pull ghcr.io/accel-sim/accel-sim-framework:ubuntu-24.04-cuda-12.8\n\n# Pull code repo\ngit clone git@github.com:accel-sim/gpgpu-sim_distribution.git\ncd gpgpu-sim_distribution\n\n# Run the image in interactive mode with gpgpu-sim mounted to `/accel-sim/gpgpu-sim_distribution` inside container\ndocker run -it --name gpgpusim -v ./:/accel-sim/gpgpu-sim_distribution\n```\n\nIf you are using [VSCode](https://code.visualstudio.com/) or [Codespace](https://github.com/features/codespaces), you setup environment with [devcontainer](https://code.visualstudio.com/docs/devcontainers/containers).\n\n- For VSCode, refer to [this guide](https://code.visualstudio.com/docs/devcontainers/containers) to setup devcontainer.\n- For Codespace, you can use this link: [https://codespaces.new/gpgpu-sim/gpgpu-sim_distribution](https://codespaces.new/gpgpu-sim/gpgpu-sim_distribution) to setup with devcontainer.\n\nTo build the simulator, you first need to configure how you want it to be built. From the root directory of the simulator, type the following commands in a bash shell (you can check you are using a bash shell by running the command \"echo $SHELL\", which should print \"/bin/bash\"):\n\n```\nsource setup_environment <build_type>\n```\n\nreplace <build_type> with `debug` or `release`. Use `release` if you need faster\nsimulation and `debug` if you need to run the simulator in gdb. If nothing is\nspecified, `release` will be used by default.\n\nNote: specifying `build_type` has no impact with the CMake build flow as it relies on the CMake variable `CMAKE_BUILD_TYPE` to determine whether to build for `release` or `debug`\n\nTo build with `cmake`, simply run the following commands:\n\n```\n# Create a release build for CMake\ncmake -B build\n\n# Or you can specify a debug build \ncmake -DCMAKE_BUILD_TYPE=Debug -B build\n\n# Build with 8 processes\ncmake --build build -j8\n\n# Install built .so to lib/ folder\ncmake --install build\n```\n\nTo build the simulator with `make`, just run\n\n```\nmake\n```\n\nAfter make is done, the simulator would be ready to use. To clean the build, run\n\n```\nmake clean\n```\n\nTo build the doxygen generated documentations, run\n\n```\nmake docs\n```\n\nTo clean the docs run\n\n```\nmake cleandocs\n```\n\nThe documentation resides at doc/doxygen/html.\n\nTo run Pytorch applications with the simulator, install the modified Pytorch library as well by following instructions [here](https://github.com/gpgpu-sim/pytorch-gpgpu-sim).\n\nBefore we run, we need to make sure the application's executable file is dynamically linked to CUDA runtime library. This can be done during compilation of your program by introducing the nvcc flag \"-lcudart\" in makefile (quotes should be excluded).\n\nTo confirm the same, type the follwoing command:\n\n`ldd <your_application_name>`\n\nYou should see that your application is using libcudart.so file in GPGPUSim directory. If the application is a Pytorch application, `<your_application_name>` should be `$PYTORCH_BIN`, which should be set during the Pytorch installation.\n\nIf running applications which use cuDNN or cuBLAS:\n\n- \nModify the Makefile or the compilation command of the application to change all the dynamic links to static ones, for example: \n  - \n`-L$(CUDA_PATH)/lib64 -lcublas` to`-L$(CUDA_PATH)/lib64 -lcublas_static`\n  - \n`-L$(CUDNN_PATH)/lib64 -lcudnn` to`-L$(CUDNN_PATH)/lib64 -lcudnn_static`\n- \n- \nModify the Makefile or the compilation command such that the following flags are used by the nvcc compiler: `-gencode arch=compute_61,code=compute_61`(the number 61 refers to the SM version. You would need to set it based on the GPGPU-Sim config `-gpgpu-ptx-force-max-capability` you use)\n\nCopy the contents of configs/QuadroFX5800/ or configs/GTX480/ to your application's working directory. These files configure the microarchitecture models to resemble the respective GPGPU architectures.\n\nTo use ptxplus (native ISA) change the following options in the configuration file to \"1\" (Note: you need CUDA version 4.0) as follows:\n\n```\n-gpgpu_ptx_use_cuobjdump 1\n-gpgpu_ptx_convert_to_ptxplus 1\n```\n\nNow To run a CUDA application on the simulator, simply execute\n\n```\nsource setup_environment <build_type>\n```\n\nUse the same <build_type> you used while building the simulator. Then just\nlaunch the executable as you would if it was to run on the hardware. By\nrunning `source setup_environment <build_type>` you change your LD_LIBRARY_PATH\nto point to GPGPU-Sim's instead of CUDA or OpenCL runtime so that you do NOT\nneed to re-compile your application simply to run it on GPGPU-Sim.\n\nTo revert back to running on the hardware, remove GPGPU-Sim from your LD_LIBRARY_PATH environment variable.\n\nThe following GPGPU-Sim configuration options are used to enable AccelWattch\n\n```\n-power_simulation_enabled 1 (1=Enabled, 0=Not enabled)\n-power_simulation_mode 0 (0=AccelWattch_SASS_SIM or AccelWattch_PTX_SIM, 1=AccelWattch_SASS_HW, 2=AccelWattch_SASS_HYBRID)\n-accelwattch_xml_file <filename>.xml\n```\n\nThe AccelWattch XML configuration file name is set to accelwattch_sass_sim.xml by default and is\ncurrently provided for SM7_QV100, SM7_TITANV, SM75_RTX2060_S, and SM6_TITANX.\nNote that all these AccelWattch XML configuration files are tuned only for SM7_QV100. Please refer to\n[https://github.com/VijayKandiah/accel-sim-framework#accelwattch-overview](https://github.com/VijayKandiah/accel-sim-framework#accelwattch-overview) for more information.\n\nRunning OpenCL applications is identical to running CUDA applications. However, OpenCL applications need to communicate with the NVIDIA driver in order to build OpenCL at runtime. GPGPU-Sim supports offloading this compilation to a remote machine. The hostname of this machine can be specified using the environment variable OPENCL_REMOTE_GPU_HOST. This variable should also be set through the setup_environment script. If you are offloading to a remote machine, you might want to setup passwordless ssh login to that machine in order to avoid having too retype your password for every execution of an OpenCL application.\n\nIf you need to run the set of applications in the NVIDIA CUDA SDK code samples then you will need to download, install and build the SDK.\n\nThe CUDA applications from the ISPASS 2009 paper mentioned above are distributed separately on github under the repo ispass2009-benchmarks. The README.ISPASS-2009 file distributed with the benchmarks now contains updated instructions for running the benchmarks on GPGPU-Sim v3.x.\n\nIf you have made modifications to the simulator and wish to incorporate new features/bugfixes from subsequent releases the following instructions may help. They are meant only as a starting point and only recommended for users comfortable with using source control who have experience modifying and debugging GPGPU-Sim.\n\nWARNING: Before following the procedure below, back up your modifications to GPGPU-Sim. The following procedure may cause you to lose all your changes. In general, merging code changes can require manual intervention and even in the case where a merge proceeds automatically it may introduce errors. If many edits have been made the merge process can be a painful manual process. Hence, you will almost certainly want to have a copy of your code as it existed before you followed the procedure below in case you need to start over again. You will need to consult the documentation for git in addition to these instructions in the case of any complications.\n\nSTOP. BACK UP YOUR CHANGES BEFORE PROCEEDING. YOU HAVE BEEN WARNED. TWICE.\n\nTo update GPGPU-Sim you need git to be installed on your system. Below we assume that you ran the following command to get the source code of GPGPU-Sim:\n\n```\n  git clone git://dev.ece.ubc.ca/gpgpu-sim\n```\n\nSince running the above command you have made local changes and we have published changes to GPGPU-Sim on the above git server. You have looked at the changes we made, looking at both the new CHANGES file and probably even the source code differences. You decide you want to incorporate our changes into your modified version of GPGPU-Sim.\n\nBefore updating your source code, we recommend you remove any object files:\n\n```\n  make clean\n```\n\nThen, run the following command in the root directory of GPGPU-Sim:\n\n```\n  git pull\n```\n\nWhile git is pulling the latest changes, conflicts might arise due to changes that you made that conflict with the latest updates. In this case, you need to resolved those conflicts manually. You can either edit the conflicting files directly using your favorite text editor, or you can use the following command to open a graphical merge tool to do the merge:\n\n```\n  git mergetool\n```\n\nNow you should test that the merged version \"works\". This means following the\nsteps for building GPGPU-Sim in the *new* README file (not this version) since\nthey may have changed. Assuming the code compiles without errors/warnings the\nnext step is to do some regression testing. At UBC we have an extensive set of\nregression tests we run against our internal development branch when we make\nchanges. In the future we may make this set of regression tests publically\navailable. For now, you will want to compile the merged code and re-run all of\nthe applications you care about (implying these applications worked for you\nbefore you did the merge). You want to do this before making further changes to\nidentify any compile time or runtime errors that occur due to the code merging\nprocess.\n\nSome applications take several hours to execute on GPGPUSim. This is because the simulator has to dump the PTX, analyze them and get resource usage statistics. This can be avoided everytime we execute the program in the following way:\n\n1. \nExecute the program by enabling “-save_embedded_ptx 1” in config file, execute the code and let cuobjdump command dump all necessary files. After this process, you will get 2 new files namely: *cuobjdump_complete_output* <some_random_name> and _1.ptx\n2. \nCreate new environment variables or include the below in your .bashrc file: \n  1. export PTX_SIM_USE_PTX_FILE=_1.ptx\n  2. export PTX_SIM_KERNELFILE=_1.ptx\n  3. export CUOBJDUMP_SIM_FILE=*cuobjdump_complete_output* <some_random_name>\n3. \nDisable -save_embedded_ptx flag, execute the code again. This will skip the dumping by cuobjdump and directly goes to executing the program thus saving time.\n\nCredits: Tor M Aamodt\n\nTo debug failing GPGPU-Sim regression tests you need to run them locally.  The fastest way to do this, assuming you are working with GPGPU-Sim versions more recent than the GPGPU-Sim dev branch circa March 28, 2018 (commit hash 2221d208a745a098a60b0d24c05007e92aaba092), is to install Docker.  The instructions below were tested with Docker CE version 18.03 on Ubuntu and Mac OS.  Docker will enable you to run the same set of regressions used by GPGPU-Sim when submitting a pull request to [https://github.com/gpgpu-sim/gpgpu-sim_distribution](https://github.com/gpgpu-sim/gpgpu-sim_distribution) and also allow you to log in and launch GPGPU-Sim in gdb so you can inspect failures.\n\n1. \nInstall Docker. On Ubuntu 14.04 and 16.04 the following instructions work: [https://docs.docker.com/install/linux/docker-ce/ubuntu/#uninstall-old-versions](https://docs.docker.com/install/linux/docker-ce/ubuntu/#uninstall-old-versions)\n2. \nClone GPGPU-Sim from your fork of GPGPU-Sim. For example: git clone [https://github.com/](https://github.com/) /gpgpu-sim_distribution.git\n3. \nRun the following command (this is all one line) to run the regressions in docker: \n\n```\ndocker run --privileged -v `pwd`:/home/runner/gpgpu-sim_distribution:rw aamodt/gpgpu-sim_regress:latest /bin/bash -c \"./start_torque.sh; chown -R runner /home/runner/gpgpu-sim_distribution; su - runner -c 'source /home/runner/gpgpu-sim_distribution/setup_environment && make -j -C /home/runner/gpgpu-sim_distribution && cd /home/runner/gpgpu-sim_simulations/ && git pull && /home/runner/gpgpu-sim_simulations/util/job_launching/run_simulations.py -c /home/runner/gpgpu-sim_simulations/util/job_launching/regression_recipies/rodinia_2.0-ft/configs.gtx1080ti.yml -N regress && /home/runner/gpgpu-sim_simulations/util/job_launching/monitor_func_test.py -v -N regress'; tail -f /dev/null\"\n```\n\n Explanation: The last part of this command, \"tail -f /dev/null\" will keep the docker container running after the regressions finish. This enables you to log into the container to run the same tests inside gdb so you can debug. The \"--privileged\" part enables you to use breakpoints inside gdb in a container. The \"-v\" part maps the current directory (with the GPGPU-Sim source code you want to test) into the container. The string \"aamodt/gpgpu-sim_regress:latest\" is a tag for a container setup to run regressions which will be downloaded from docker hub. The portion starting with /bin/bash is a set of commands run inside a bash shell inside the container. E.g., the command start_torque.sh starts up a queue manager inside the container. If the above command stops with the message \"fatal: unable to access ' [https://github.com/tgrogers/gpgpu-sim_simulations.git/](https://github.com/tgrogers/gpgpu-sim_simulations.git/) ': Could not resolve host: github.com\" this likely means your computer sits behind a firewall which is blocking access to Google's name servers (e.g., 8.8.8.8).  To get around this you will need to modify th above command to point to your local DNS server.  Lookup your DNS server IP address which we will call <DNS_IP_ADDRESS> below.  On Ubuntu run \"ifconfig\" to lookup the network interface connecting your computer to the network.  Then run \"nmcli device show \" to find the IP address of your DNS server.  Modify the above command to include \"--dns <DNS_IP_ADDRESS>\" after \"run\", E.g.,\n\n```\ndocker run --dns <DNS_IP_ADDRESS> --privileged -v `pwd`:/home/runner/gpgpu-sim_distribution:rw aamodt/gpgpu-sim_regress:latest /bin/bash -c \"./start_torque.sh; chown -R runner /home/runner/gpgpu-sim_distribution; su - runner -c 'source /home/runner/gpgpu-sim_distribution/setup_environment && make -j -C /home/runner/gpgpu-sim_distribution && cd /home/runner/gpgpu-sim_simulations/ && git pull && /home/runner/gpgpu-sim_simulations/util/job_launching/run_simulations.py -c /home/runner/gpgpu-sim_simulations/util/job_launching/regression_recipies/rodinia_2.0-ft/configs.gtx1080ti.yml -N regress && /home/runner/gpgpu-sim_simulations/util/job_launching/monitor_func_test.py -v -N regress'; tail -f /dev/null\"\n```\n\n4. \nFind the CONTAINER ID associated with your docker container by running \"docker ps\".\n5. \nLog into the container by running the command: \n\n```\ndocker exec -it <CONTAINER_ID> /bin/bash -c \"su -l runner\"`\n```\n\n The container is running Ubuntu 16.04 and has screen, cscope and vim installed (if you find a favorite Linux tool missing, it is fairly easy to create derived containers that have additional tools).\n6. \nLookup the directory of the regression test you want to debug by going to the regression log file directory: \n\n```\ncd /home/runner/gpgpu-sim_simulations/util/job_launching/logfiles\n```\n\n7. \nThe file \"failed_job_log_sim_log.regress..txt\" includes information about the failed test including its simulation directory. For the following example, I'll assume the first failing test was \"hotspot-rodinia-2.0-ft-30_6_40___data_result_30_6_40_txt--GTX1080Ti\" for which the simulation directory is /home/runner/gpgpu-sim_simulations/util/job_launching/../../sim_run_4.2/hotspot-rodinia-2.0-ft/30_6_40___data_result_30_6_40_txt/GTX1080Ti/\n8. \nChange to the simulation directory using: \n\n```\ncd <simulation_directory>\n```\n\n E.g., `cd /home/runner/gpgpu-sim_simulations/util/job_launching/../../sim_run_4.2/hotspot-rodinia-2.0-ft/30_6_40___data_result_30_6_40_txt/GTX1080Ti/`This directory should contain a file called \"torque.sim\" that contains commands used to launch the simulation during regression tests. We will modify this file to enable us to re-run the regression test in gdb. This directory should also contain a file containing the standard output during the regression test. This file will end in .o where is the torque queue manager job number. For the running example for me this file is called \"hotspot-rodinia-2.0-ft-30_6_40___data_result_30_6_40_txt.o2\". Open this file to determine the LD_LIBRARY_PATH settings used when launching the simulation. Look for a line that starts \"doing: export LD_LIBRARY_PATH\" and copy the entire line starting with \"export LD_LIBRARY_PATH ...\"\n9. \nPaste the \"export LD_LIBRARY_PATH ...\" line into the bash shell to set LD_LIBRARY_PATH. E.g., \n\n```\nexport LD_LIBRARY_PATH=/home/runner/gpgpu-sim_simulations/util/job_launching/../../sim_run_4.2/gpgpu-sim-builds/libcudart_gpgpu-sim_git-commit-177d02254ae38b6331b17dd6cd139b570a03c589_modified_0.so:/gpgpu-sim/usr/local/gcc-4.5.4/lib64:/gpgpu-sim/usr/local/gcc-4.5.4/lib:/gpgpu-sim/usr/local/gcc-4.5.4/lib/gcc/x86_64-unknown-linux-gnu/lib64/:/gpgpu-sim/usr/local/gcc-4.5.4/lib/gcc/x86_64-unknown-linux-gnu/4.5.4/:/usr/lib/x86_64-linux-gnu:/home/runner/gpgpu-sim_distribution/lib/gcc-4.5.4/cuda-4020/release:/gpgpu-sim/usr/local/gcc-4.5.4/lib64:/gpgpu-sim/usr/local/gcc-4.5.4/lib:/gpgpu-sim/usr/local/gcc-4.5.4/lib/gcc/x86_64-unknown-linux-gnu/lib64/:/gpgpu-sim/usr/local/gcc-4.5.4/lib/gcc/x86_64-unknown-linux-gnu/4.5.4/:/usr/lib/x86_64-linux-gnu:\n```\n\n10. \nIn the same shell, build the debug version of GPGPU-Sim then return to the directory above: \n\n```\npushd ~/gpgpu-sim_distribution/\nsource setup_environment debug\nmake\npopd\n```\n\n11. \nOpen and edit torque.sim and preface the very last line with \"gdb --args \". After editing the last line in torque.sim should look something like: \n\n```\ngdb --args /home/runner/gpgpu-sim_simulations/util/job_launching/../../benchmarks/bin/4.2/release/hotspot-rodinia-2.0-ft 30 6 40 ./data/result_30_6_40.txt\n```\n\n12. \nRe-run the regression test in gdb by sourcing the torque.sim file: \n\n```\n. torque.sim\n```\n\n This will put you in at the (gdb) prompt. Setup any breakpoints needed and run.\n\nThe `gpu->is_SST_mode()` conditionals in the codebase address architectural differences between GPGPU-Sim's original design and SST integration, primarily focusing on two areas:\n\n- SST-Specific Hardware Configuration\n  - Cache bypass: SST mode intercepts interconnect packets to redirect them externally instead of using GPGPU-Sim's native cache system.\n  - Component initialization: Guards hardware setup steps that only apply to SST's simulation environment.\n- Frontend-Backend Coupling\n  - In standard GPGPU-Sim:\n    - CUDA frontend (separate thread) asynchronously pushes operations to stream manager\n    - Backend simulator consumes operations independently via `cycle()` calls\n  - In SST mode:\n    - Single-threaded execution requires synchronization callbacks between Balar's frontend event handlers and backend clock ticks\n    - Prevents deadlocks where:\n      - Blocking CUDA requests wait for stream clearance\n      - Stream processing depends on backend `cycle()` advancement\n      - SST progression halts until current event handling completes\n- In standard GPGPU-Sim:\n\nThis coupling necessitated modified stream management to replace GPGPU-Sim's native busy-wait approach with SST-compatible synchronization triggers.\n\nFor detailed documentation on SST-integration, checkout the [SST elements documentation](https://sst-simulator.org/sst-docs/docs/elements/intro).", "url": "https://wpnews.pro/news/gpgpu-sim-cycle-level-simulator-for-nvidia-gpus-with-cuda-or-opencl-workloads", "canonical_source": "https://github.com/gpgpu-sim/gpgpu-sim_distribution", "published_at": "2026-10-09 22:27:23+00:00", "updated_at": "2026-10-09 22:55:01.501309+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "developer-tools", "machine-learning"], "entities": ["GPGPU-Sim", "Nvidia", "CUDA", "OpenCL", "Accel-Sim", "AccelWattch", "NVBit", "AerialVision"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gpgpu-sim-cycle-level-simulator-for-nvidia-gpus-with-cuda-or-opencl-workloads", "markdown": "https://wpnews.pro/news/gpgpu-sim-cycle-level-simulator-for-nvidia-gpus-with-cuda-or-opencl-workloads.md", "text": "https://wpnews.pro/news/gpgpu-sim-cycle-level-simulator-for-nvidia-gpus-with-cuda-or-opencl-workloads.txt", "jsonld": "https://wpnews.pro/news/gpgpu-sim-cycle-level-simulator-for-nvidia-gpus-with-cuda-or-opencl-workloads.jsonld"}}