# Extreme RAM consumption of an md simulation

**URL:** <https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248>\
**Category:** User discussions\
**Tags:** mdrun, simulation-setup\
**Created:** [September 22, 2023, 4:06pm UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248 "2023-09-22T16:06:01Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![max.m90](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@max.m90](https://gromacs.bioexcel.eu/u/max.m90)\
**Post date:** [September 22, 2023, 4:06pm UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/1 "2023-09-22T16:06:02Z")

</div>

GROMACS version: 2022.5  
GROMACS modification: No  
I encountered an ‘OUT\_OF\_MEMORY’ error with the following job details for an md simulation with ~250,000 particles:

Nodes: 16  
Cores per node: 64  
CPU Utilized: 310-17:42:30  
CPU Efficiency: 97.49% of 318-17:50:56 core-walltime  
Job Wall-clock time: 07:28:14  
Memory Utilized: 5.31 TB (estimated maximum)  
Memory Efficiency: 392.86% of 1.35 TB (86.43 GB/node)

srun -n 512 -c 2 gmx\_mpi mdrun -ntomp 2 -deffnm md

I’m interested in understanding why the simulation required such an excessive amount of RAM and what steps I can take to optimize memory usage in future GROMACS simulations on this cluster.

The workflow (solvation, ions, energy minimization,…) until the md run is pretty similar to the gromacs tutorial.

Any insights or recommendations would be greatly appreciated.

Thanks in advance!

Max

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [September 25, 2023, 7:52am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/2 "2023-09-25T07:52:30Z")

</div>

We can’t say much without more information about your setup. But it is strange that you can run energy minimization but not MD. What changes did you make in the mdp parameters between EM and MD?

---

<div class="post-metadata">

**Author:** ![max.m90](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@max.m90](https://gromacs.bioexcel.eu/u/max.m90)\
**Post date:** [September 25, 2023, 8:43am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/3 "2023-09-25T08:43:19Z")

</div>

Thank you for your answer!

Here is the setup:

- I am investigating a protein
- ff: amber99sb-star-ildn-q-tip4pd.ff
- As solvent model I used tip4p in a dodecahedron box (-c -d 1.0 -bt dodecahedron)
- EM.mdp: is the same as in the lysozyme in water tutotrial  
emtol = 1000.0  
emstep = 0.01  
nsteps = 50000
- MD.mdp changes:  
nsteps = 500000000 (1μsecond)  
dt = 0.002  
continuation = no  
en\_vel = yes   
gen\_seed = -1   
For the rest of the MD.mdp file I used the default values

CPU: Intel Gold 6130 S2/C16/T2

In our group we also had this problem with different simulation as soon as 16 nodes are to be claimed.  
With 8 nodes the simulation works but the performance is not satisfying.

I hope I could share some useful information.  
Thanks for the help

---

<div class="post-metadata">

**Author:** ![MagnusL](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/magnusl/32/2380_2.png) [@MagnusL](https://gromacs.bioexcel.eu/u/MagnusL)\
**Post date:** [September 25, 2023, 8:53am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/4 "2023-09-25T08:53:31Z")

</div>

Does it make a difference if you run fewer tasks per node?  
E.g. `srun -n 64 -c 16 gmx_mpi mdrun -ntomp 16 -deffnm md`

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [September 25, 2023, 9:06am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/5 "2023-09-25T09:06:06Z")

</div>

It strange that it works on 8 but not on 16 nodes. Using more nodes reduces the memory requirement per node.

What is the memory usage on 8 nodes?

---

<div class="post-metadata">

**Author:** ![max.m90](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@max.m90](https://gromacs.bioexcel.eu/u/max.m90)\
**Post date:** [September 25, 2023, 9:59am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/6 "2023-09-25T09:59:53Z")

</div>

@hess

The MaxRSS on 8 nodes were 87148K.

@MagnusL

I haven´t done it yet, but thats something I will try.

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [September 25, 2023, 10:38am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/7 "2023-09-25T10:38:38Z")

</div>

87 MB can’t be correct.

---

<div class="post-metadata">

**Author:** ![max.m90](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@max.m90](https://gromacs.bioexcel.eu/u/max.m90)\
**Post date:** [September 25, 2023, 10:58am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/8 "2023-09-25T10:58:08Z")

</div>

@hess

Yes that was a mistake sorry.  
that should be right

Nodes: 8  
Cores per node: 64  
CPU Utilized: 2027-09:21:42  
CPU Efficiency: 98.99% of 2048-01:25:20 core-walltime  
Job Wall-clock time: 4-00:00:10  
Memory Utilized: 21.28 GB (estimated maximum)  
Memory Efficiency: 3.08% of 691.41 GB (86.43 GB/node)

---

<div class="post-metadata">

**Author:** ![hess](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/hess/32/416_2.png) [@hess](https://gromacs.bioexcel.eu/u/hess)\
**Post date:** [September 25, 2023, 1:52pm UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/9 "2023-09-25T13:52:43Z")

</div>

So the memory usage goes from 21 GB to 5.3 TB when doubling the number of nodes. Or do I misunderstand something? Such a change is very unlikely to come from GROMACS. It could be some hidden bug that suddenly triggers order of magnitude more memory usage wen doubling the number of nodes.

What are the last few lines in the log file of the run that goes out of memory?

---

<div class="post-metadata">

**Author:** ![max.m90](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@max.m90](https://gromacs.bioexcel.eu/u/max.m90)\
**Post date:** [September 26, 2023, 8:53am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/10 "2023-09-26T08:53:36Z")

</div>

Yes exactly. The problem is that I don´t have an idea where it is comming from.  
I already talked to our cluster support, but their only answer was to double the memory capacity for the next run but that wonn´t help if we are talking about TB.

I dont have the log files anymore, but I am currently waiting for my SLURM job to get started with the adapted command line MagnusL provided.  
If I encounter the same problem again I can provide the log file.

Thanks for the input!!

---

<div class="post-metadata">

**Author:** ![max.m90](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@max.m90](https://gromacs.bioexcel.eu/u/max.m90)\
**Post date:** [September 28, 2023, 7:13am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/11 "2023-09-28T07:13:27Z")

</div>

Hey @MagnusL

the simulation with your provided command is running smoothly wie 90ns/day.  
Thank you for your advice. But still I do not really understand the problem, can you explain me on which grounds you decided to change the values of -n, -c and -ntomp?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![MagnusL](https://dub1.discourse-cdn.com/flex017/user_avatar/gromacs.bioexcel.eu/magnusl/32/2380_2.png) [@MagnusL](https://gromacs.bioexcel.eu/u/MagnusL)\
**Post date:** [September 28, 2023, 8:00am UTC](https://gromacs.bioexcel.eu/t/extreme-ram-consumption-of-an-md-simulation/7248/12 "2023-09-28T08:00:23Z")

</div>

Good to hear that it helped at least.

Unfortunately I don’t have any specific grounds for that recommendation. It’s just that my personal experience has shown that when running on more than 3 or 4 nodes it has been more efficient not to increase the total number of MPI tasks. I.e., to lower the number of tasks per node. I haven’t had the reported RAM issues, though, so it was just that I thought that 512 MPI tasks sounded a bit high.

Edit: It’s quite possible that `srun -n 128 -c 8 gmx_mpi mdrun -ntomp 8 -deffnm md` would be worth trying, perhaps also `srun -n 256 -c 4 gmx_mpi mdrun -ntomp 4 -deffnm md`. You might see a difference in performance.
