Leading Unnatural Amino Acids & PEG Derivatives CDMO | TideChem

Cookie Settings

We and our affiliates use cookie technology to provide you with customized content that interests you, identify visitors, ensure secure login, and collect data. Click “Accept All” to accept all cookies and jump directly to the website.

Accept All
About
Drug Development and Regulatory Studies
Home / About / Drug Development and Regulatory Studies

Peptide Sequences: A Practical Guide

2026-08-28 Posted by TideChem view:94

A peptide sequence is the exact order of amino acid residues in a peptide. This order determines many of the peptide’s properties, including molecular weight, charge, hydrophobicity, solubility, conformation, biological activity, and susceptibility to enzymatic degradation.

Peptide sequences are normally written from the N-terminus to the C-terminus. However, residue order alone may not completely define a peptide. Terminal groups, stereochemistry, chemical modifications, cyclization, and disulfide connectivity may also need to be specified.

How Are Peptide Sequences Written?

Peptide sequences can be written using three-letter or one-letter amino acid codes.

For example:

Three-letter format: Gly–Ala–Ser–Lys
One-letter format: GASK

The 20 standard amino acids and their codes are:

Amino acid Three-letter One-letter
Alanine Ala A
Arginine Arg R
Asparagine Asn N
Aspartic acid Asp D
Cysteine Cys C
Glutamine Gln Q
Glutamic acid Glu E
Glycine Gly G
Histidine His H
Isoleucine Ile I
Leucine Leu L
Lysine Lys K
Methionine Met M
Phenylalanine Phe F
Proline Pro P
Serine Ser S
Threonine Thr T
Tryptophan Trp W
Tyrosine Tyr Y
Valine Val V

One-letter codes are convenient for databases and long sequences. Three-letter codes are often clearer for short peptides, D-amino acids, non-natural residues, and modified amino acids.

N-Terminus and C-Terminus

Peptides have a defined direction:

N-terminus → C-terminus

The N-terminus contains the first residue, while the C-terminus contains the final residue. Unless otherwise stated, sequences should always be written in this direction.

Conventional solid-phase peptide synthesis builds the chain from the C-terminal residue toward the N-terminus, but the completed peptide is still reported from N to C.

Writing a sequence backward produces a different peptide with potentially different structure and biological activity.

Sequence and Composition Are Not the Same

Amino acid composition describes which amino acids are present and how many of each are contained in the peptide. A peptide sequence describes their exact order.

For example, GASK and KSAG have the same amino acid composition and molecular weight, but they are different peptides.

This distinction is important because intact mass and amino acid analysis may confirm composition without proving residue order. Sequence confirmation generally requires tandem mass spectrometry, Edman degradation, or comparison with a verified reference.

How Sequence Affects Peptide Properties

Charge and Solubility

Lysine and arginine usually contribute positive charge, while aspartate and glutamate usually contribute negative charge. Histidine may change protonation state near physiological pH.

The number and distribution of charged residues influence aqueous solubility, chromatographic behavior, membrane interaction, and receptor binding.

Hydrophobicity

Leucine, isoleucine, valine, phenylalanine, tryptophan, and other nonpolar residues increase hydrophobicity.

Hydrophobic residues can strengthen target or membrane binding, but excessive hydrophobicity may cause poor solubility, aggregation, or nonspecific adsorption.

Conformation

Amino acid order influences whether a peptide favors an α-helix, β-structure, turn, or flexible conformation.

Proline restricts backbone geometry, while glycine increases flexibility. Cyclization and disulfide bonds can further constrain the peptide structure.

Proteolytic Stability

Proteases recognize particular sequences and conformations. Replacing one residue may remove a cleavage site, but it may also reduce target affinity.

Stability can sometimes be improved with:

  • D-amino acids
  • N-methyl amino acids
  • Non-natural residues
  • Cyclization
  • Terminal modification

These changes must be evaluated experimentally because an improvement in stability may be accompanied by reduced activity or solubility.

Terminal and Side-Chain Modifications

Two peptides with the same residue order can behave differently when their terminal groups or side chains are modified.

Common N-terminal modifications include:

  • Acetylation
  • Formylation
  • Biotinylation
  • Fluorescent labeling
  • Lipidation

Common C-terminal forms include:

  • Free carboxylic acid
  • Amide
  • Ester
  • Alcohol
  • Aldehyde

Possible side-chain modifications include phosphorylation, methylation, acetylation, glycosylation, sulfation, lipidation, and stable-isotope labeling.

Each modification should be identified by its exact position. A description such as “phosphorylated peptide” is incomplete when several residues could be phosphorylated.

Linear, Cyclic, and Disulfide-Linked Peptides

A linear peptide contains one continuous chain with an N-terminus and C-terminus.

Cyclic peptides may be connected through:

  • Head-to-tail cyclization
  • Side-chain lactam bridges
  • Disulfide bonds
  • Thioether linkages
  • Hydrocarbon staples

For peptides containing several cysteines, the sequence alone does not define which cysteines are connected. Disulfide connectivity should be stated separately.

Branched and multichain peptides also require a clear description of each chain and every connection point.

How Are Peptide Sequences Confirmed?

Mass Spectrometry

Intact-mass analysis can determine whether the observed molecular weight agrees with the expected peptide.

However, correct mass does not always prove correct sequence. Peptides containing the same residues in different orders may have identical molecular weights.

Tandem Mass Spectrometry

MS/MS fragments the peptide backbone and provides information about residue order and modification sites. Fragment coverage around critical positions should be reviewed carefully.

Leucine and isoleucine have the same exact mass and are difficult to distinguish using routine mass spectrometry. Additional sequence information or specialized fragmentation may be required.Leucine/isoleucine MS study

Edman Degradation

Edman degradation identifies amino acids sequentially from a free N-terminus. It can confirm short sequences but is less suitable for peptides with blocked N-termini or complex mixtures.Edman sequencing method

HPLC or UPLC

Chromatography is used primarily to assess purity. It does not independently confirm peptide sequence.

A high HPLC purity value should not be interpreted as proof of correct identity, net peptide content, or biological activity.

Sequence-Related Impurities

Repeated coupling and deprotection reactions during peptide synthesis can produce:

  • Deletion sequences
  • Truncated peptides
  • Epimerized residues
  • Oxidized products
  • Incomplete deprotection products
  • Incorrect disulfide isomers
  • Side-chain reaction products

Long, hydrophobic, aggregation-prone, or extensively modified sequences are generally more difficult to synthesize and purify.

Analytical characterization should therefore be selected according to sequence complexity and intended use.

Designing a Peptide Sequence

A practical design process should consider:

  1. Biological purpose: Define whether the peptide is an antigen, enzyme substrate, receptor ligand, inhibitor, analytical standard, or therapeutic lead.
  2. Correct biological region: Confirm species, isoform, residue numbering, sequence version, and mature-chain boundaries.
  3. Solubility: Review charge, hydrophobicity, concentration, and proposed solvent.
  4. Stability: Identify protease-sensitive, oxidation-prone, or deamidation-prone residues.
  5. Synthesis feasibility: Consider length, hydrophobic segments, modifications, cyclization, and disulfide formation.
  6. Controls: Include suitable unmodified, scrambled, substituted, or unrelated control peptides.

A scrambled sequence is not automatically a valid negative control. It should retain relevant physicochemical properties without preserving the biological recognition motif.

Information Needed for Custom Peptide Synthesis

A complete peptide request should specify:

  • Sequence from N to C
  • L- or D-configuration
  • Non-natural amino acids
  • N-terminal modification
  • C-terminal form
  • Modification sites
  • Cyclization method
  • Disulfide connectivity
  • Required quantity and purity
  • Counterion preference
  • Analytical requirements
  • Net peptide content requirement
  • Endotoxin or sterility requirements, when applicable

A request stating only the sequence and “95% purity” is incomplete. HPLC purity does not define identity, peptide content, counterion, water content, or biological activity.

Frequently Asked Questions

What is a peptide sequence?

It is the exact order of amino acid residues in a peptide, written from the N-terminus to the C-terminus.

Can two peptides have the same molecular weight but different sequences?

Yes. Peptides with the same amino acid composition in different orders may have identical molecular weights.

How is a peptide sequence verified?

Common methods include intact-mass analysis, tandem mass spectrometry, Edman degradation, amino acid analysis, and comparison with reference standards.

Does HPLC purity confirm the sequence?

No. HPLC measures chromatographic purity. Mass spectrometry or another identity method is needed to support sequence confirmation.

Why must terminal modifications be specified?

N-terminal acetylation and C-terminal amidation change peptide charge, molecular weight, stability, and potentially biological activity.

What does X mean in a sequence?

X usually represents an unknown or unspecified residue. It should be resolved before custom peptide synthesis.

Conclusion

Peptide sequences define the order of amino acid residues and strongly influence molecular weight, charge, solubility, conformation, stability, and biological activity.

A complete peptide specification should also include terminal forms, stereochemistry, modifications, cyclization, and disulfide connectivity. Reliable characterization requires separate evaluation of identity, purity, content, structure, and activity.

References

  1. IUPAC-IUB Nomenclature for Amino Acids and Peptides
  2. UniProt Sequence Annotation Guidance
  3. N-Terminal Sequence Analysis by Edman Chemistry
  4. De Novo Peptide Sequencing by Mass Spectrometry
  5. MS Differentiation of Leucine and Isoleucine

Hot Articles

Categories