# Why Fatal error: Unexpected cudaStreamQuery failure happend in gromacs2019?

**URL:** <https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238>\
**Category:** User discussions\
**Tags:** mdrun\
**Created:** [September 21, 2023, 9:02am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238 "2023-09-21T09:02:17Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![ladimafakher](https://avatars.discourse-cdn.com/v4/letter/l/cc9497/32.png) [@ladimafakher](https://gromacs.bioexcel.eu/u/ladimafakher)\
**Post date:** [September 21, 2023, 9:02am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/1 "2023-09-21T09:02:17Z")

</div>

GROMACS version:2019  
GROMACS modification: Yes/No no  
I have Gromacs 2019. when I want to run mdrun by GPU command I faced this error " Fatal error: Unexpected cudaStreamQuery failure: an illegal memory access was encountered. what is the problem and how I can solve it?

---

<div class="post-metadata">

**Author:** ![MagnusL](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/magnusl/32/2380_2.png) [@MagnusL](https://gromacs.bioexcel.eu/u/MagnusL)\
**Post date:** [September 21, 2023, 10:04am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/2 "2023-09-21T10:04:21Z")

</div>

The version that you are using is more than four years old and is no longer actively supported. Is there a reason why you are not using the 2023 version (2023.2 is the latest patch release)? If you have to stick to the 2019 version, make sure that you are using the latest patched version (2019.6). Hopefully that will fix the problem.

---

<div class="post-metadata">

**Author:** ![ladimafakher](https://avatars.discourse-cdn.com/v4/letter/l/cc9497/32.png) [@ladimafakher](https://gromacs.bioexcel.eu/u/ladimafakher)\
**Post date:** [September 21, 2023, 10:55am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/3 "2023-09-21T10:55:55Z")

</div>

thanks for your answer

---

<div class="post-metadata">

**Author:** ![Yel21](https://avatars.discourse-cdn.com/v4/letter/y/b2d939/32.png) [@Yel21](https://gromacs.bioexcel.eu/u/Yel21)\
**Post date:** [October 24, 2023, 1:27pm UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/4 "2023-10-24T13:27:25Z")

</div>

Hello, @MagnusL !

I had the same problem with Gromacs 2023.1 with CUDA 12.2 and nvidia-driver-535 on Ubuntu 23.10.

```auto
Program: gmx mdrun, version 2023.1
Source file: src/gromacs/gpu_utils/cudautils.cuh (line 190)

Fatal error:
Unexpected cudaStreamQuery failure. CUDA error #719 (cudaErrorLaunchFailure):
unspecified launch failure.

```

The simulations were performed on a machine based on 1 GPU RTX4090 and an Intel i9-13900. I tried to simulate different systems on the machine using various Gromacs versions from 2021 to 2023.1 with different CUDA toolkits 11.7, 12.1, and 12.2; in all cases, the same problem occurred.

What do you think? How can I solve this problem?

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [October 25, 2023, 8:56am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/5 "2023-10-25T08:56:22Z")

</div>

Does the same tpr run stably when running on CPU only?

---

<div class="post-metadata">

**Author:** ![Yel21](https://avatars.discourse-cdn.com/v4/letter/y/b2d939/32.png) [@Yel21](https://gromacs.bioexcel.eu/u/Yel21)\
**Post date:** [October 25, 2023, 11:33am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/6 "2023-10-25T11:33:54Z")

</div>

Hi, @hess !

I did not try to start calculation only using CPU. I will write here what happen in a couple of days.

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [October 25, 2023, 12:56pm UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/7 "2023-10-25T12:56:30Z")

</div>

Do you get the error at step 0? If so, then trying on a CPU takes a few minutes, even on a laptop.

---

<div class="post-metadata">

**Author:** ![Yel21](https://avatars.discourse-cdn.com/v4/letter/y/b2d939/32.png) [@Yel21](https://gromacs.bioexcel.eu/u/Yel21)\
**Post date:** [October 26, 2023, 8:30am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/8 "2023-10-26T08:30:25Z")

</div>

Hello, @hess ,

> [@hess](#):
>
> Does the same tpr run stably when running on CPU only?

When I started the calculation only with the CPU using the command

```auto
gmx mdrun -v -deffnm 300 -bonded cpu -pme cpu -nb cpu -s topol.tpr -cpi 300.cpt

```

The speed is much slower, and the calculation is stable for 200 ns. I did not perform the simulation for longer.

> [@hess](#):
>
> Do you get the error at step 0? If so, then trying on a CPU takes a few minutes, even on a laptop.

This error occurs 100–150 ns after the simulation begins. Yes, I initially tried to start the simulation using the CPU and then switched to the GPU. However, the error did not disappear.

I also tried changing the screen lock time, as suggested here [Fatal error: Unexpected cudaStreamQuery failure](https://gromacs.bioexcel.eu/t/fatal-error-unexpected-cudastreamquery-failure/899) . However, in my case, it did not help me get rid with error.

I believed it might be related to the hardware of the machine, but until now, I have not figured out whether it is related to the videocard, RAM, or CPU.

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [October 27, 2023, 7:13am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/9 "2023-10-27T07:13:26Z")

</div>

If it occurs after so much time that makes it difficult to debug and also difficult to guess where it comes from. If there is a bug in GROMACS the chance is very low that it would only pop up after 100 million time steps. It could of course be that you system becomes unstable and that the first error that is triggered is a CUDA error. But then it is also unlikely that your simulation becomes unstable only after so many steps.

---

<div class="post-metadata">

**Author:** ![Wai](https://avatars.discourse-cdn.com/v4/letter/w/bcef8e/32.png) [@Wai](https://gromacs.bioexcel.eu/u/Wai)\
**Post date:** [January 19, 2024, 3:16am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/10 "2024-01-19T03:16:36Z")

</div>

Dear all,

Are there any updates on this? I’ve got this too.  
I used the “lysozyme in water” tutorial to test my gromacs and the error only appeared in the middle of the md stage (extended to 100 ns), at step 18100000 (time 36200 ps).  
I have one RTX 4090, Gromacs 2023.2, cuda 12.2, nvidia-driver-535.154.05, ubuntu 20.04 LTS.

```auto
------------------------------------------------------
Program: gmx mdrun, version 2023.2
Source file: src/gromacs/gpu_utils/device_stream.cu (line 100)
Function: DeviceStream::synchronize() const::<lambda()>

Assertion failed:
Condition: stat == cudaSuccess
cudaStreamSynchronize failed. CUDA error #719 (cudaErrorLaunchFailure):
unspecified launch failure.

For more information and tips for troubleshooting, please check the GROMACS
website at http://www.gromacs.org/Documentation/Errors
-------------------------------------------------------

```

#mdrun  
#gpu

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [January 22, 2024, 11:13am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/11 "2024-01-22T11:13:44Z")

</div>

No, no updates.

I have no clue if this would be due to a bug in GROMACS, an instability in your simulation (which still should not result in a assertion failure, but an error message) or a CUDA driver issue.

Has this happened only once or multiple times?

---

<div class="post-metadata">

**Author:** ![Wai](https://avatars.discourse-cdn.com/v4/letter/w/bcef8e/32.png) [@Wai](https://gromacs.bioexcel.eu/u/Wai)\
**Post date:** [January 23, 2024, 1:50am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/12 "2024-01-23T01:50:30Z")

</div>

Thanks for responding. I have just set this system up. So far, 6-7 trials have been run on the lysozyme in water md all failed at some point in the middle. I have not been able to complete this task at all. We tried another system that is a small peptide in water a few times (6-7), which also always failed with this error at some point. I also tried 2023.3 with no luck. At the moment, I cannot use this system to do any MD work at all.

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [January 23, 2024, 7:39am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/13 "2024-01-23T07:39:44Z")

</div>

The only known way to get such a failure is with an unstable system and PME on GPU, then an illegal memory access can occur when particles are far outside of the unit cell. Could you run with the option -pme cpu to check if this might avoid the issue?

---

<div class="post-metadata">

**Author:** ![Yel21](https://avatars.discourse-cdn.com/v4/letter/y/b2d939/32.png) [@Yel21](https://gromacs.bioexcel.eu/u/Yel21)\
**Post date:** [January 24, 2024, 7:52am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/14 "2024-01-24T07:52:37Z")

</div>

Hello @hess,

I have conducted experiments on several systems at varying temperature conditions. The outcome is contingent upon the stability of the systems. In some instances, switching to the `-pme cpu` mode and carrying out further simulations for several nanoseconds may facilitate a return to the simulation using the `-pme gpu` mode. But it is not work for each case.

---

<div class="post-metadata">

**Author:** ![Wai](https://avatars.discourse-cdn.com/v4/letter/w/bcef8e/32.png) [@Wai](https://gromacs.bioexcel.eu/u/Wai)\
**Post date:** [January 24, 2024, 9:50am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/16 "2024-01-24T09:50:04Z")

</div>

Hi, I tried with “-pme cpu” and got an error and core dump:

Step 4882651 Pressure scaling more than 1%. This may mean your system is not yet equilibrated. Use of Parrinello-Rahman pressure coupling during equilibration can lead to simulation instability, and is discouraged.  
Segmentation fault (core dumped)

Does it reveal anything? I ran this same task (without “-pme cpu”) on a different server with 4 GPU and it completed without any issues.

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [January 24, 2024, 10:08am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/17 "2024-01-24T10:08:27Z")

</div>

Then I guess that the CUDA error is triggered by instability of the simulation. The question is then why your simulation becomes unstable.

What are the mdp settings for your thermostat and barostat?

---

<div class="post-metadata">

**Author:** ![Wai](https://avatars.discourse-cdn.com/v4/letter/w/bcef8e/32.png) [@Wai](https://gromacs.bioexcel.eu/u/Wai)\
**Post date:** [January 25, 2024, 2:02am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/18 "2024-01-25T02:02:53Z")

</div>

Hi,

As follows:

```auto
; Temperature coupling is on
tcoupl = V-rescale ; modified Berendsen thermostat
tc-grps = Protein Non-Protein ; two coupling groups - more acc
urate
tau_t = 0.1 0.1 ; time constant, in ps
ref_t = 300 300 ; reference temperature, one for
 each group, in K
; Pressure coupling is on
pcoupl = Parrinello-Rahman ; Pressure coupling on in NPT
pcoupltype = isotropic ; uniform scaling of box vectors
tau_p = 2.0 ; time constant, in ps
ref_p = 1.0 ; reference pressure, in bar
compressibility = 4.5e-5 ; isothermal compressibility of 
water, bar^-1

```

The test was taken from the lysozyme tutorial. The only change I made was to run it for longer (change from 1 ns to 100 ns).

[http://www.mdtutorials.com/gmx/lysozyme/08\_MD.html](http://www.mdtutorials.com/gmx/lysozyme/08_MD.html)

---

<div class="post-metadata">

**Author:** ![MagnusL](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/magnusl/32/2380_2.png) [@MagnusL](https://gromacs.bioexcel.eu/u/MagnusL)\
**Post date:** [January 25, 2024, 7:19am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/19 "2024-01-25T07:19:36Z")

</div>

Thanks. It is quite probable that such an old tutorial was never tested (or meant for) running 100 ns and that there are instabilities that show up when you do.

It would be interesting to see if your simulation is stable if you modify the thermostat and barostat to this:

```auto
; Temperature coupling is on
tcoupl = V-rescale
tc-grps = Protein Non-Protein ; two coupling groups - more acc
urate
tau_t = 1 1 ; time constant, in ps
ref_t = 300 300 ; reference temperature, one for
each group, in K
; Pressure coupling is on
pcoupl = c-rescale ; Pressure coupling on in NPT
pcoupltype = isotropic ; uniform scaling of box vectors
tau_p = 5.0 ; time constant, in ps
ref_p = 1.0 ; reference pressure, in bar
compressibility = 4.5e-5 ; isothermal compressibility of
water, bar^-1

```

I.e., change the barostat to c-rescale (more stable - there is a risk of oscillations when using Parrinello-Rahman, especially with low tau-p) and change tau-t and tau-p.

If this does not help, it is possible that the system is not equilibrated enough to run this long. The NVT and NPT stages might need to be extended as well.

---

<div class="post-metadata">

**Author:** ![Wai](https://avatars.discourse-cdn.com/v4/letter/w/bcef8e/32.png) [@Wai](https://gromacs.bioexcel.eu/u/Wai)\
**Post date:** [January 25, 2024, 8:25am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/20 "2024-01-25T08:25:31Z")

</div>

Thank you for the suggestion. I tried that which resulted in the same error message.

```auto
-------------------------------------------------------
Program: gmx mdrun, version 2023.3
Source file: src/gromacs/gpu_utils/device_stream.cu (line 100)
Function: DeviceStream::synchronize() const::<lambda()>

Assertion failed:
Condition: stat == cudaSuccess
cudaStreamSynchronize failed. CUDA error #719 (cudaErrorLaunchFailure):
unspecified launch failure.

For more information and tips for troubleshooting, please check the GROMACS
website at http://www.gromacs.org/Documentation/Errors
-------------------------------------------------------

```

In fact, I tried the same tutorial on a different system (but different CPU and GPU) with similar software (Ubuntu 20.04LTS, Gromacs 2023.3, cuda 12.2) and it has no problem completing the 100ns MD. I had a feeling that the real cause may be hardware related…

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [January 25, 2024, 8:29am UTC](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238/21 "2024-01-25T08:29:03Z")

</div>

Try with tau\_p=5. Maybe we’re lucky that the (too) short period for the barostat is causing the issues. Or switch to C-rescale, which is anyhow what we would recommend.

[Next page](https://gromacs.bioexcel.eu/t/why-fatal-error-unexpected-cudastreamquery-failure-happend-in-gromacs2019/7238.md?page=2)
