(a) Pairwise comparisons of in silico (blue) R e and in cell (green) FRET efficiencies. In silico points are the mean end-to-end distance (R ee ) of 3 converged simulation repeats, and error bars are their standard deviation. In cell markers and error bars are the average and standard deviation of the medians of all biological repeats (n = listed in Fig. 2a). Significance is determined by an unpaired two-sided Student’s t-test without multiple testing correction, where the asterisks associated with each sequence pair denote the resulting p-values: * p < 0.05, ** p < 0.01, *** p < 0.001, **** p < 0.0001. Sequences with evenly distributed net-positive charges all increase in ensemble dimension as the fraction of charged residues increases. (b) For the same sequence parameters but with clustered charges, in silico sequences display no significant ensemble change, while in cell sequences expand. This may imply environmental conditions influencing ensemble behaviour in cells. (c) Schematic of how GOOSE was used to generate sequences tested. All sequences here are 60 residues. Five sequences are generated with specified parameters; R e is predicted using ALBATROSS for each sequence, and the mean is reported as the condition average. This is repeated, titrating along both the number of prolines and charged residues, until the heatmaps in panels d and e are complete. (d) Heatmap of average R e as proline content and the number of charged residues increase, with evenly distributed charges. Here, as both the number of prolines and the number of charged residues increase, R e increases. Compared sequences for which proline content differs in Fig. 2f are marked as labelled stars. (e) Heatmap for sequences with clustered charges. Here, there is no consistent dependence on the number of charged residues, but increases in proline content expand the ensemble. (f) Same format as panels a and b, but with predicted R e instead of simulated R e . In cell E f markers and error bars remain the average and standard deviation of the medians of all biological repeats (n = listed in Fig. 2a). Predicted R e markers are the mean of the condition average, while the error bars are the standard deviation (n = 5 sequences). Significance is determined by an unpaired two-sided Student’s t-test without multiple testing correction, where the asterisks associated with each sequence pair denote the resulting p-values: * p < 0.05, ** p < 0.01, *** p < 0.001, **** p < 0.0001. For evenly distributed charges, all ensembles expand as FCR increases and/or as proline counts increase. (g) When charges are clustered, increasing FCR and proline count concomitantly (left) expands the ensemble under both predicted and in-cell conditions, whereas when only FCR is significantly increased (right), the ensemble compacts in cell but shows no predicted change.
Rational design of disordered proteins for sequence–function investigation
Why This Matters
This study highlights how sequence design influences the conformational behavior of disordered proteins, revealing environmental effects within cells that differ from in silico predictions. Understanding these dynamics is crucial for developing targeted therapeutics and bioengineering applications involving intrinsically disordered proteins. The research underscores the importance of considering cellular context when predicting protein behavior, advancing the precision of protein design.
Key Takeaways
- Sequence charge distribution impacts protein ensemble dimensions differently in vitro and in vivo.
- Proline content and charge clustering influence the conformational flexibility of disordered proteins.
- Integrating computational predictions with cellular data enhances understanding of protein behavior in biological environments.
Get alerts for these topics