# Removing water molecules from tpr file while conserving correct protein chain assignment

**URL:** <https://gromacs.bioexcel.eu/t/removing-water-molecules-from-tpr-file-while-conserving-correct-protein-chain-assignment/12327>\
**Category:** User discussions\
**Created:** [July 7, 2025, 12:51am UTC](https://gromacs.bioexcel.eu/t/removing-water-molecules-from-tpr-file-while-conserving-correct-protein-chain-assignment/12327 "2025-07-07T00:51:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![mverissi](https://avatars.discourse-cdn.com/v4/letter/m/e99b99/32.png) [@mverissi](https://gromacs.bioexcel.eu/u/mverissi)\
**Post date:** [July 7, 2025, 12:51am UTC](https://gromacs.bioexcel.eu/t/removing-water-molecules-from-tpr-file-while-conserving-correct-protein-chain-assignment/12327/1 "2025-07-07T00:51:45Z")

</div>

GROMACS version: 2025.1  
GROMACS modification: No

Dear all,

I have performed four sets of MD simulations of a globular protein interacting with a coiled coil with 30,000 frames each, aiming at studying which residues interact more frequently at the interface between the protein and the coil. Since the coiled coil is very long, I have ~ 12K protein atoms and ~750K water atoms. I intend to use MDAnalysis for this, since I find it easier to extract more detailed information with it.

In MDAnalysis, I have to load the .tpr and .xtc files, select the atoms and run the analysis. However, the frames are loaded one at a time and I don’t need the waters. One solution would be to remove the solvent waters, or even have a trajectory with only the interface region. This is easy with gmx trjconv.

However, when I try to convert the tpr file, the information on chains is lost, and MDAnalysis reads all chains as “Protein in water”. Even creating an index file with the different chains didn’t work.

Another solution would be to create a coordinate file with the charges of each atom, but I didn’t find an option for creating such a file with gmx trjconv.

Would anyone know how to create the tpr file with the correct chain information when removing the water molecules?

Best,

Marcos Verissimo Alves  
Universidade Federal Fluminense, Brazil

---

<div class="post-metadata">

**Author:** ![BjarneF](https://avatars.discourse-cdn.com/v4/letter/b/9d8465/32.png) [@BjarneF](https://gromacs.bioexcel.eu/u/BjarneF)\
**Post date:** [July 9, 2025, 8:55am UTC](https://gromacs.bioexcel.eu/t/removing-water-molecules-from-tpr-file-while-conserving-correct-protein-chain-assignment/12327/2 "2025-07-09T08:55:47Z")

</div>

Hi Marcos,

I have been in this situation before and solved it by loading a .top topology instead of a .tpr file. Since .top is plain text, it’s much easier to modify - in your case, it should be as simple as deleting the water entry from the [molecules] record.  
MDAnalysis will by default assume your .top is an AMBER topology file, so you will have to pass the appropriate keyword to mda.Universe:

u = mda.Universe(topol.top, traj.xtc, topology\_format=“ITP”)

This should preserve the segment ID information.

See here for further reference, for example for inclusion of files from non-standard directories: [5.10. ITP topology parser — MDAnalysis 2.0.0 documentation](https://docs.mdanalysis.org/2.0.0/documentation_pages/topology/ITPParser.html)

Best wishes,  
Bjarne

---

<div class="post-metadata">

**Author:** ![mverissi](https://avatars.discourse-cdn.com/v4/letter/m/e99b99/32.png) [@mverissi](https://gromacs.bioexcel.eu/u/mverissi)\
**Post date:** [July 9, 2025, 5:30pm UTC](https://gromacs.bioexcel.eu/t/removing-water-molecules-from-tpr-file-while-conserving-correct-protein-chain-assignment/12327/3 "2025-07-09T17:30:24Z")

</div>

Hi Bjarne,

Thanks a lot for the tip! I’ll look into it.

Best,

Marcos
