# A series of performance benchmarks for MD Apps, including GROMACS

**URL:** <https://gromacs.bioexcel.eu/t/a-series-of-performance-benchmarks-for-md-apps-including-gromacs/7078>\
**Category:** User discussions\
**Tags:** mdrun, gpu, mdrun-performance\
**Created:** [August 24, 2023, 11:04pm UTC](https://gromacs.bioexcel.eu/t/a-series-of-performance-benchmarks-for-md-apps-including-gromacs/7078 "2023-08-24T23:04:41Z")\
**Posts on this page:** 1\
**Showing post:** 14

<div class="post-metadata">

**Author:** ![al42and](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/al42and/32/1393_2.png) [@al42and](https://gromacs.bioexcel.eu/u/al42and)\
**Post date:** [August 30, 2023, 6:42pm UTC](https://gromacs.bioexcel.eu/t/a-series-of-performance-benchmarks-for-md-apps-including-gromacs/7078/14 "2023-08-30T18:42:10Z")

</div>

First, thank you for such a thorough analysis! And for taking your time to translate it to English!

> [@Entropy\_YU](#):
>
> However, this hardly works on the gfx1100, as it gets stuck after running dozens of tests and then the OS crashes. I’ve used different gfx1100 (7900XTX) GPUs and other hardwares but got the same results. Therefore, it took me a lot of time to measure the GROMACS mdrun performance on gfx1100.

Have you tried setting `HIPSYCL_RT_MAX_CACHED_NODES=0` environment variable, as described in [GROMACS get stuck AMD GPU](https://gromacs.bioexcel.eu/t/gromacs-get-stuck-amd-gpu/6625)?

We still don’t have a good understanding of the root cause of the problem, but we suspect that the problem might be caused by the hipSYCL caching behavior, where it submits tasks to the GPU in bursts, with is handled poorly by the AMD HSA runtime sometimes. Setting `HIPSYCL_RT_MAX_CACHED_NODES=0` will force immediate submission avoiding this potential problem. With the latest hipSYCL, is almost always improves performance too, but mostly for small systems.

> [@Entropy\_YU](#):
>
> However, I’ve noticed that a lot of new verified commits have been made after July 25, and I’m looking forward to the potential performance improvements!

I don’t think there were many performance-related changes since then. They added a new backend (OpenCL) and a new programming model (C++ Standard Parallelism / stdpar), which are significant changes, but neither is used by GROMACS.

* * *

Speaking of Intel A770: one can get slightly better performance when using Double-batched FFT library instead of MKL. For A770, [double-batched FFT](https://github.com/intel/double-batched-fft-library) should be compiled with `-DCMAKE_CXX_COMPILER=icpx -DCMAKE_BUILD_TYPE=Release -DBUILD_SHARED_LIBS=YES -DNO_DOUBLE_PRECISION=ON`, then [the GROMACS should be told to use it](https://manual.gromacs.org/documentation/current/install-guide/index.html#using-double-batched-fft-library).

In my tests, the speed-up is around ~10% on STMV when using oneAPI 2023.2 (same as yours), so nothing dramatic. Just FYI.

---

_[View the full topic](https://gromacs.bioexcel.eu/t/a-series-of-performance-benchmarks-for-md-apps-including-gromacs/7078)._
