Questions and Answers

The server predicts plausible 3D models of peptide aggregates formed by amyloid-prone peptides that assemble into in- register, parallel β-sheets. To perform a prediction, the amino acid sequence of the amyloidogenic peptide must be provided. Computational tasks can be submitted using either Basic or Advanced mode.

Basic mode:

  • Provide the amino acid sequence of the aggregating peptide (capital letters only)
  • All calculations are performed using default parameter values.
  • Supported sequence length: 10–100 residues.
  • After submission, you will receive a link to the results page (if an email address is provided).
  • A typical run takes from a few hours up to a few days, depending on sequence length and server load.
  • Once the job is finished, the results can be viewed on the results page (link provided during submission or sent via email, if available).

Advanced mode:

  • Advanced mode lets you modify simulation settings, including: temperature range for Replica Exchange Monte Carlo (REMC); secondary structure definition; pH value; number of Monte Carlo (MC) cycles;
  • Optionally, additional restraints can be applied to mimic inter-chain disulfide (S–S) bonds between cysteine residues.

For an example job (input and output), see: (link do examples).

No. The InRegPep assembler does not predict amyloid propensity. It is designed to generate plausible 3D aggregate models for peptides that are already considered amyloid-prone. Each simulation returns 10 different aggregate conformations (the Top10 model set). Benchmarking indicated that the Top10 set included models resembling experimental structures for 10 of the 15 analyzed systems (RMSD < 5 Å) (link to benchmark).

In Advanced mode, you can change:

  • Temperature range for replicas in the REMC algorithm (default: 2.0-1.0).
  • Number of Monte Carlo steps (default: 50).
  • Secondary structure (C - coil, E – extended; default: C).
  • pH (default: 7.0).

You can also add distance restraints between pairs residues to mimic inter-chain disulphide bonds (see section Examples).

Note: The default settings were selected to balance accuracy and reasonable run time.

When the submitted job is complete, the results are available via the link generated during submission. Results include:

  • The set of Top10 best-scored aggregate models, each composed of 5 interacting peptides
  • A stability label for each model: “Plausible” or “Unstable”
  • Automated analyses for each model, including:
    • CABSenergy vs. pcaRMSD plot
    • RMSF per residue plot
    • β-sheet content per residue plot
    • pcaRMSD value (measure of translational symmetry of peptides forming aggregate)
    • Normalized contact metrics (measure of contacts between layers)

For models ranked Plausible, the server also generates a short protofilament model consisting of 15 layers.

All outputs including structures can be downloaded as a .zip archive.

The archive also includes traj.zip - a compressed PDB-format trajectory of representative aggregate models in Cα-trace representation obtained from all docking trajectories. In particular, it includes the lowest-energy structure, the structure with the highest symmetry (as indicated by the pcaRMSD parameter), and 10 structures identified using k-means clustering for each Monte Carlo run. The total number of unique structures is equal to 12 × the number of Monte Carlo runs.

All results are stored on the server for at least 14 days after job completion by default. Please download all data before this time frame, as the results may be deleted afterward.

No. The InRegPep assembler is designed only to model aggregates that form in-register, parallel β-sheets.

Depending on the length of the submitted peptide sequence (i.e., system size) and the current server load, a job may take from a few hours up to a few days.

During aggregate prediction, a large number of models is generated using coarse-grained docking with the CABS multichain model. From these, 20 models are selected based on a combination of energy-based criteria, structural symmetry (quantified by the pcaRMSD parameter), and hierarchical clustering. The selected models are then reconstructed into all-atom representations and refined by energy minimization. Next, the refined aggregates are subjected to molecular dynamics (MD) simulations in explicit solvent. Each MD trajectory is analyzed, and stability is assessed based on β-sheet content, inter-chain contact patterns, and structural symmetry. Models that remain stable and maintain ordered packing are labeled Plausible, whereas models that show low translational symmetry, reduced β-sheet content, or unfavorable contact patterns are labeled Unstable.

The CABS (C-Alpha, Beta, Side-chain) program is a well-established CG protein modeling tool for predicting protein dynamics and protein structures at coarse-graining levels. In the CABS model, the majority of single amino acids are represented by two, three, or four united atoms. These pseudo atoms (also called united residue units) are placed in the C-alpha positions, the center of the C-alpha-C-alpha pseudo-bonds, the C-beta positions, and the center of the remaining fragment of the side-chain of heavy atoms (when appropriate). C-alpha positions are restricted to the nodes of the underlying lattice network. This assumption requires small fluctuations in the C-alpha-C-alpha distances. The positions of the remaining pseudo-atoms are not limited to the discrete lattice patterns and are defined by the C-alpha chains. The centers of side chains represent the most probable positions observed in protein fragments (based on a statistical analysis of the protein structures from the Protein Data Bank) with specific geometry of the main chain.

Such a design has several implications that distinguish CABS from other CG models of similar resolution. Firstly, since the C-alpha traces can be stored as a string of integer coordinates (defining positions of other pseudo-atoms), most short-range distances (and related interactions) can be stored in large data tables, allowing rapid computations of local movement and interactions. This way, CABS-based simulations are noticeably faster than simulations of continuous models of comparable resolution. The grid space of the main chain discretization is sufficiently dense to avoid directional biases, which is the typical weakness of crude lattice models of protein structures, commonly used by the polymer/protein physics community in the past. The inherent resolution of the CABS representation is approximately 0.5 Å for the main-chain atoms, and 1.5 Å for the entire protein, due to the predefined positions of the side chain centers. Side-chain packing is not perfect; however, the “single conformers” restriction can be relaxed by small fluctuations of the main chain traces.

The CABS force field comprises knowledge-based statistical potentials derived from structural regularities that characterize known protein structures. Statistical potentials that control CABS include hardcore (excluded volume) of the main chain pseudo atoms and C-beta atoms, soft repulsive cores of side chains centers, and contact energy for the specific distances between side chains that depend on their mutual orientation. The directional model of the main chain hydrogen bonds also has a character of contact interactions with the contact distances adjusted to the distances between the appropriate united atoms in adjacent beta-strands and within alpha-helices. There is also a multi-body potential controlling preference for specific geometries of the main chains, mainly right-handed helical fragments and extended fragments of strands present in beta-sheets. Predicted secondary structure preferences can modify these main chain geometrical preferences. The effects of the surrounding solvent are treated in an implicit, highly simplified manner by appropriate shifts of the contact potentials and a small (almost negligible) centrosymmetric attractive force located in the centers of mass of single-domain proteins. In the CABS-dock applications of CABS, the centrosymmetric potential can be ignored.

The Monte Carlo dynamics scheme controls sampling of the CABS conformational space. Local moves are randomly selected and attempt to modify small fragments of the modeled chain. These random moves are accepted or rejected according to the asymmetric Metropolis criterion. Long sequences of local pseudo-random motions produce a trajectory that provides a long-time dynamics picture of the modeled system. Of course, as in all CG simulations, the time unit of resulting trajectories is poorly defined and needs to be re-adjusted. It may be achieved by comparing a short fragment of CG trajectories with long MD simulations, producing conformational modifications on a similar scale.

For more information, please refer to the following literature:

  1. Kolinski, A., Protein modeling and structure prediction with a reduced representation. Acta Biochimica Polonica, 2004. 51(2): p. 349-371.
  2. Kurcinski, M., A. Badaczewska‐Dawid, M. Kolinski, A. Kolinski, and S. Kmiecik, Flexible docking of peptides to proteins using CABS‐dock. Protein Science, 2020. 29(1): p. 211-222.
  3. Puławski, W., A. Koliński, and M. Koliński, Integrative modeling of diverse protein-peptide systems using CABS-dock. PLOS Computational Biology, 2023. 19(7): p. e1011275.
  4. Puławski, W., A. Koliński, and M. Koliński, Multiscale modeling of protofilament structures: A case study on insulin amyloid aggregates. International Journal of Biological Macromolecules, 2025. 285: p. 138382.
  5. Kmiecik, S., D. Gront, M. Kolinski, L. Wieteska, A.E. Dawid, and A. Kolinski, Coarse-grained protein models and their applications. Chemical reviews, 2016. 116(14): p. 7898-7936.

Use the Top 10 models as a set of alternative hypotheses for the aggregate structure. A typical interpretation guide:

  • Plausible vs. Unstable: Start with models labeled Plausible as the most reliable candidates.
  • CABS-energy vs. pcaRMSD plot: Shows the correlation between model energy and the translational symmetry of interacting peptides in the CABS docking trajectory.
  • RMSF per residue: Lower RMSF values indicate more rigid and ordered regions of the aggregate, while higher RMSF values suggest flexible segments.
  • β-sheet content per residue: Highlights which residues most consistently form β-sheet-rich structures across all the models in the MD trajectory.
  • Contact metrics: Help determine whether the assembly exhibits regular packing.

For predicted aggregate structures labeled Plausible, a protofilament model consisting of 15 layers of interacting peptides is generated.

Yes. The method generates multiple aggregate conformations for the same amyloid-prone peptide, which may help explore structural polymorphism.

All generated models are included in archive - ZIP file that can be downloaded from the Results page after the job is completed. In addition, the archive file also includes traj.zip — a compressed PDB- format trajectory of representative aggregate models in Cα-trace representation obtained from all docking trajectories. In particular, it contains the lowest-energy structure, the structure with the highest symmetry (as indicated by the pcaRMSD parameter), and 10 structures identified using k-means clustering for each docking run. The total number of unique structures is equal to 12 × the number of docking runs.

When using or referring to results obtained with InRegPep, please cite:

  • “InRegPep: a web server for modeling of parallel beta-sheet fibrillar aggregates of amyloidogenic peptides” (manuscript in preparation)

For questions or issues, please contact:

  • Wojciech Puławski, PhD; e-mail: wpulawski@imdik.pan.pl
  • Szymon Niewieczerzał, PhD; email: sz.niewieczerzal@cent.uw.edu.pl
  • Michał Koliński, PhD DSc; e-mail: mkolinski@imdik.pan.pl