The server predicts plausible 3D models of peptide aggregates formed by amyloid-prone peptides that assemble into in- register, parallel β-sheets. To perform a prediction, the amino acid sequence of the amyloidogenic peptide must be provided. Computational tasks can be submitted using either Basic or Advanced mode.
Basic mode:
Advanced mode:
For an example job (input and output), see: (link do examples).
In Advanced mode, you can change:
You can also add distance restraints between pairs residues to mimic inter-chain disulphide bonds (see section Examples).
Note: The default settings were selected to balance accuracy and reasonable run time.
When the submitted job is complete, the results are available via the link generated during submission. Results include:
For models ranked Plausible, the server also generates a short protofilament model consisting of 15 layers.
All outputs including structures can be downloaded as a .zip archive.
The archive also includes traj.zip - a compressed PDB-format trajectory of representative aggregate models in Cα-trace representation obtained from all docking trajectories. In particular, it includes the lowest-energy structure, the structure with the highest symmetry (as indicated by the pcaRMSD parameter), and 10 structures identified using k-means clustering for each Monte Carlo run. The total number of unique structures is equal to 12 × the number of Monte Carlo runs.
All results are stored on the server for at least 14 days after job completion by default. Please download all data before this time frame, as the results may be deleted afterward.
No. The InRegPep assembler is designed only to model aggregates that form in-register, parallel β-sheets.
Depending on the length of the submitted peptide sequence (i.e., system size) and the current server load, a job may take from a few hours up to a few days.
During aggregate prediction, a large number of models is generated using coarse-grained docking with the CABS multichain model. From these, 20 models are selected based on a combination of energy-based criteria, structural symmetry (quantified by the pcaRMSD parameter), and hierarchical clustering. The selected models are then reconstructed into all-atom representations and refined by energy minimization. Next, the refined aggregates are subjected to molecular dynamics (MD) simulations in explicit solvent. Each MD trajectory is analyzed, and stability is assessed based on β-sheet content, inter-chain contact patterns, and structural symmetry. Models that remain stable and maintain ordered packing are labeled Plausible, whereas models that show low translational symmetry, reduced β-sheet content, or unfavorable contact patterns are labeled Unstable.
The CABS (C-Alpha, Beta, Side-chain) program is a well-established CG protein modeling tool for predicting protein dynamics and protein structures at coarse-graining levels. In the CABS model, the majority of single amino acids are represented by two, three, or four united atoms. These pseudo atoms (also called united residue units) are placed in the C-alpha positions, the center of the C-alpha-C-alpha pseudo-bonds, the C-beta positions, and the center of the remaining fragment of the side-chain of heavy atoms (when appropriate). C-alpha positions are restricted to the nodes of the underlying lattice network. This assumption requires small fluctuations in the C-alpha-C-alpha distances. The positions of the remaining pseudo-atoms are not limited to the discrete lattice patterns and are defined by the C-alpha chains. The centers of side chains represent the most probable positions observed in protein fragments (based on a statistical analysis of the protein structures from the Protein Data Bank) with specific geometry of the main chain.
Such a design has several implications that distinguish CABS from other CG models of similar resolution. Firstly, since the C-alpha traces can be stored as a string of integer coordinates (defining positions of other pseudo-atoms), most short-range distances (and related interactions) can be stored in large data tables, allowing rapid computations of local movement and interactions. This way, CABS-based simulations are noticeably faster than simulations of continuous models of comparable resolution. The grid space of the main chain discretization is sufficiently dense to avoid directional biases, which is the typical weakness of crude lattice models of protein structures, commonly used by the polymer/protein physics community in the past. The inherent resolution of the CABS representation is approximately 0.5 Å for the main-chain atoms, and 1.5 Å for the entire protein, due to the predefined positions of the side chain centers. Side-chain packing is not perfect; however, the “single conformers” restriction can be relaxed by small fluctuations of the main chain traces.
The CABS force field comprises knowledge-based statistical potentials derived from structural regularities that characterize known protein structures. Statistical potentials that control CABS include hardcore (excluded volume) of the main chain pseudo atoms and C-beta atoms, soft repulsive cores of side chains centers, and contact energy for the specific distances between side chains that depend on their mutual orientation. The directional model of the main chain hydrogen bonds also has a character of contact interactions with the contact distances adjusted to the distances between the appropriate united atoms in adjacent beta-strands and within alpha-helices. There is also a multi-body potential controlling preference for specific geometries of the main chains, mainly right-handed helical fragments and extended fragments of strands present in beta-sheets. Predicted secondary structure preferences can modify these main chain geometrical preferences. The effects of the surrounding solvent are treated in an implicit, highly simplified manner by appropriate shifts of the contact potentials and a small (almost negligible) centrosymmetric attractive force located in the centers of mass of single-domain proteins. In the CABS-dock applications of CABS, the centrosymmetric potential can be ignored.
The Monte Carlo dynamics scheme controls sampling of the CABS conformational space. Local moves are randomly selected and attempt to modify small fragments of the modeled chain. These random moves are accepted or rejected according to the asymmetric Metropolis criterion. Long sequences of local pseudo-random motions produce a trajectory that provides a long-time dynamics picture of the modeled system. Of course, as in all CG simulations, the time unit of resulting trajectories is poorly defined and needs to be re-adjusted. It may be achieved by comparing a short fragment of CG trajectories with long MD simulations, producing conformational modifications on a similar scale.
For more information, please refer to the following literature:
Use the Top 10 models as a set of alternative hypotheses for the aggregate structure. A typical interpretation guide:
For predicted aggregate structures labeled Plausible, a protofilament model consisting of 15 layers of interacting peptides is generated.
Yes. The method generates multiple aggregate conformations for the same amyloid-prone peptide, which may help explore structural polymorphism.
All generated models are included in archive - ZIP file that can be downloaded from the Results page after the job is completed. In addition, the archive file also includes traj.zip — a compressed PDB- format trajectory of representative aggregate models in Cα-trace representation obtained from all docking trajectories. In particular, it contains the lowest-energy structure, the structure with the highest symmetry (as indicated by the pcaRMSD parameter), and 10 structures identified using k-means clustering for each docking run. The total number of unique structures is equal to 12 × the number of docking runs.
When using or referring to results obtained with InRegPep, please cite:
For questions or issues, please contact: