# GROMACS-SYCL Build with CUDA Support

**URL:** <https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195>\
**Category:** User discussions\
**Created:** [May 20, 2021, 3:08pm UTC](https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195 "2021-05-20T15:08:45Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![vdle](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/vdle/32/844_2.png) [@vdle](https://gromacs.bioexcel.eu/u/vdle)\
**Post date:** [May 20, 2021, 3:08pm UTC](https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195/1 "2021-05-20T15:08:45Z")

</div>

GROMACS version:2021-sycl  
GROMACS modification: No

Hi,  
I’m trying to builld the sycl version of gromacs using intel-llvm compiler with cuda backend.  
The cmake options are:  
$ cmake …  
-DGMX\_GPU=SYCL  
-DGMX\_BUILD\_OWN\_FFTW=ON  
-DCMAKE\_C\_COMPILER=clang  
-DCMAKE\_CXX\_COMPILER=clang++  
However, gmx binary failed to regconized V100:  
#0: name: Tesla V100-PCIE-32GB, vendor: NVIDIA Corporation, device version: 0.0, status: incompatible (please recompile with correct GMX\_OPENCL\_NB\_CLUSTER\_SIZE of 4)  
#1: name: SYCL host device, vendor: , device version: 1.2, status: incompatible  
As I understand -DGMX\_OPENCL\_NB\_CLUSTER\_SIZE=4 is reserved for Intel GPUs.

If someone has successfully build SYCL version, I appreciate some of your insights.

---

<div class="post-metadata">

**Author:** ![rschulz](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/rschulz/32/30_2.png) [@rschulz](https://gromacs.bioexcel.eu/u/rschulz)\
**Post date:** [May 20, 2021, 3:57pm UTC](https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195/2 "2021-05-20T15:57:11Z")

</div>

The 2021-sycl version doesn’t support Nvidia GPUs. You can use that version only with Intel GPUs. Work is ongoing on the master branch to add support for all GPUs and also different SYCL compilers.

If you want to try it you need to:

- start with master branch
- merge in the branch sz\_SYCL-nbnxm-local-mem-reduction (this branch is under review as [Implement generic j-reduction in nbnxm SYCL kernels (!1410) · Merge requests · GROMACS / GROMACS · GitLab](https://gitlab.com/gromacs/gromacs/-/merge_requests/1410))
- in cmake/gmxManageSYCL.cmake add -fsycl-targets=nvptx64-nvidia-cuda-sycldevice after -fsycl

Because it is still work-in-progress it isn’t optimized yet and the current performance reflects that.  
As always any user is invited to help with the effort to improve GROMACS. Let us know if you are interested.

---

<div class="post-metadata">

**Author:** ![vdle](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/vdle/32/844_2.png) [@vdle](https://gromacs.bioexcel.eu/u/vdle)\
**Post date:** [May 21, 2021, 6:40am UTC](https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195/3 "2021-05-21T06:40:30Z")

</div>

Thanks for the tips. I was able to build gromacs from master branch following your instructions, yet the error persists.  
I listed the steps here in case I overlooked something.

[git checkout]  
git clone [GROMACS / GROMACS · GitLab](https://gitlab.com/gromacs/gromacs)  
git checkout remotes/origin/sz\_SYCL-nbnxm-local-mem-reduction

[cmake/gmxManageSYCL.cmake : before]  
" “CXX” DISABLE\_SYCL\_CXX\_FLAGS SYCL\_CXX\_FLAGS “-fsycl -fsycl-device-code-split=per\_kernel”)

[cmake/gmxManageSYCL.cmake : after]  
" “CXX” DISABLE\_SYCL\_CXX\_FLAGS SYCL\_CXX\_FLAGS “-fsycl -fsycl-targets=nvptx64-nvidia-cuda-sycldevice”)  
(-fsycl-device-code-split caused configure error)

[cmake]  
cmake … (same options as the opening post)  
make

[test commands]  
gmx mdrun -noconfout -nsteps 10000 -nb gpu -s stmv.tpr -tunepme -v

Regarding performance, I strive to compare the performance of the trifecta:  
gmx-sycl (xe\_hp) vs. gmx-sycl (v100) vs. gmx-cuda (v100)

---

<div class="post-metadata">

**Author:** ![rschulz](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/rschulz/32/30_2.png) [@rschulz](https://gromacs.bioexcel.eu/u/rschulz)\
**Post date:** [May 21, 2021, 3:28pm UTC](https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195/4 "2021-05-21T15:28:43Z")

</div>

Sorry I forgot some extra steps:

- You need the cmake option -DGMX\_GPU\_NB\_CLUSTER\_SIZE=8
- You need the environment variable GMX\_GPU\_DISABLE\_COMPATIBILITY\_CHECK=1 when running mdrun

If you interest is a fair performance comparison you want to wait until it is properly working. It just started to barely work and we haven’t yet done the work to optimize for performance. We expect a lot of performance improvements over the next few months. Those will be included in the next GROMACS release with the first beta release scheduled for around September.

---

<div class="post-metadata">

**Author:** ![vdle](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/vdle/32/844_2.png) [@vdle](https://gromacs.bioexcel.eu/u/vdle)\
**Post date:** [May 22, 2021, 12:56pm UTC](https://gromacs.bioexcel.eu/t/gromacs-sycl-build-with-cuda-support/2195/5 "2021-05-22T12:56:17Z")

</div>

Thanks.

V100s were now recognized:  
Number of GPUs detected: 3  
#0: name: Tesla V100-PCIE-32GB, vendor: NVIDIA Corporation, device version: 0.0, status: compatible  
#1: name: Tesla V100-PCIE-32GB, vendor: NVIDIA Corporation, device version: 0.0, status: compatible  
#2: name: SYCL host device, vendor: , device version: 1.2, status: compatible

Unfortunately, it still crashed with the following message.  
pi\_die: cuda\_piEnqueueEventsWaitWithBarrier not implemented

This is an unresolved issue with intel-llvm:

> <https://github.com/intel/llvm/issues/3218>
>
> \*\*Describe the bug\*\*
> The PI function piEnqueueEventsWaitWithBarrier is not impl…emented by CIDA PI plugin.
> 
> \*\*To Reproduce\*\*
> Use the barrier test (currently disabled for CUDA): https://github.com/intel/llvm-test-suite/blob/intel/SYCL/Basic/enqueue\_barrier.cpp
> 
> \*\*Environment (please complete the following information):\*\*
> 
> \- OS: Windows/Linux
> \- Target device and vendor: CUDA GPU
> \- DPC++ version: any
> \- Dependencies version: n/a

I guess I have to wait until the sycl implementation matures more.
