Leading Unnatural Amino Acids & PEG Derivatives CDMO | TideChem

Cookie Settings

We and our affiliates use cookie technology to provide you with customized content that interests you, identify visitors, ensure secure login, and collect data. Click “Accept All” to accept all cookies and jump directly to the website.

Accept All
About
Drug Development and Regulatory Studies
Home / About / Drug Development and Regulatory Studies

Semaglutide Peptide Sequence Explained

2026-08-31 Posted by TideChem view:205

Semaglutide is a modified glucagon-like peptide-1 receptor agonist derived from human GLP-1(7-37). Its peptide backbone contains 31 amino acid residues, but the complete molecule cannot be described by the amino acid sequence alone. It also includes a non-natural amino acid and a fatty diacid side chain attached through a defined linker.

These structural modifications improve resistance to enzymatic degradation and promote albumin binding while retaining GLP-1 receptor activity. Understanding the complete semaglutide peptide sequence is therefore important for peptide synthesis, analytical method development, impurity identification, and pharmaceutical research.

What Is the Semaglutide Peptide Sequence?

The semaglutide peptide backbone is:

His-Aib-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys-Glu-Phe-Ile-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly-OH

The lysine residue corresponding to Lys26 in native GLP-1 carries a side chain composed of:

  • One γ-glutamic acid unit
  • Two 8-amino-3,6-dioxaoctanoic acid spacer units
  • One C18 fatty diacid

The spacer units may also be described as ADO, AEEA or OEG units, depending on the naming convention used by the supplier or analytical laboratory.

Semaglutide has the molecular formula C187H291N45O59 and a molecular weight of 4113.58 g/mol according to official product information. PubChem classifies it as both a polypeptide and a lipopeptide. PubChem, DailyMed

Understanding Semaglutide Numbering

Semaglutide is based on GLP-1(7-37), so scientific literature normally uses the residue numbering of native GLP-1 rather than numbering the semaglutide chain from 1 to 31.

Structural feature GLP-1 numbering Position in the 31-residue chain
N-terminal histidine His7 1
Aib substitution Aib8 2
Side-chain attachment Lys26 20
Arginine substitution Arg34 28
C-terminal glycine Gly37 31

This distinction explains why Aib8 is sometimes described as the second residue of semaglutide. Both statements are correct, but they use different numbering systems.

How Does Semaglutide Differ from Human GLP-1?

The human GLP-1(7-37) sequence is:

His-Ala-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys-Glu-Phe-Ile-Ala-Trp-Leu-Val-Lys-Gly-Arg-Gly-OH

Semaglutide retains 29 of these 31 amino acid positions, giving it approximately 94% sequence identity with native GLP-1. It contains two amino acid substitutions and one major side-chain modification:

  1. Ala8 is replaced by Aib8
  2. Lys34 is replaced by Arg34
  3. Lys26 is modified with a linker and C18 fatty diacid

The European Medicines Agency describes semaglutide as an Aib8, Arg34-GLP-1(7-37) analogue with a side chain attached to Lys26. EMA assessment report

Why Are These Modifications Important?

Aib8 improves enzymatic stability

Aib is 2-aminoisobutyric acid, a non-proteinogenic amino acid. It replaces alanine at position 8, which is close to the N-terminus and involved in recognition by dipeptidyl peptidase-4.

This substitution increases resistance to DPP-4-mediated degradation. Aib does not have a standard one-letter amino acid code, so a conventional one-letter sequence cannot represent semaglutide accurately without additional annotation.

Arg34 controls the acylation site

Native GLP-1 contains lysine residues at positions 26 and 34. Replacing Lys34 with arginine leaves Lys26 as the principal site for controlled side-chain attachment.

This modification helps prevent the formation of differently acylated positional variants during manufacturing.

The Lys26 side chain promotes albumin binding

The ε-amino group of Lys26 is connected to a C18 fatty diacid through a hydrophilic spacer. The spacer helps separate the lipid group from the receptor-binding peptide region, while the fatty diacid supports reversible albumin association.

Albumin binding reduces rapid renal clearance and contributes to the prolonged pharmacokinetic profile of semaglutide. The importance of the fatty acid and linker design was described in the original semaglutide discovery research. Journal of Medicinal Chemistry

Why the Peptide Sequence Alone Is Not Enough

A sequence containing 31 amino acids is not necessarily complete semaglutide. The identity of the molecule also depends on:

  • The presence and stereochemistry of Aib8
  • The Lys26 attachment site
  • The precise linker composition
  • The structure of the C18 fatty diacid
  • The free N-terminus and C-terminal carboxylic acid
  • The counterion and hydration state
  • The profile of sequence-related and conjugation-related impurities

An unmodified semaglutide backbone, sometimes described as a semaglutide main chain or des-acyl semaglutide, is an intermediate rather than the complete active molecule.

Likewise, an amino acid sequence, a research-grade peptide and an approved pharmaceutical product should not be treated as equivalent. Pharmaceutical equivalence also requires appropriate manufacturing controls, formulation, stability data, biological testing and regulatory authorization.

Semaglutide Synthesis Considerations

Regulatory documents describe commercial semaglutide production as yeast fermentation followed by chemical modification and purification. Research-scale material may also be prepared using solid-phase peptide synthesis, fragment condensation or hybrid approaches.

Important manufacturing challenges include:

  • Efficient incorporation of sterically hindered Aib
  • Orthogonal protection of the Lys26 side chain
  • Selective attachment of the linker and fatty diacid
  • Control of deletion and insertion sequences
  • Prevention of amino acid epimerization
  • Separation of closely related acylation impurities
  • Removal of residual reagents and counterions

The lipid side chain can alter chromatographic behavior and solubility. Purification methods developed for ordinary hydrophilic peptides may therefore require different gradients, stationary phases or sample-preparation conditions.

How Is the Sequence Confirmed?

No single analytical result is sufficient to establish the complete identity of semaglutide. A suitable characterization package may include:

Intact-mass analysis: LC-MS or high-resolution MS confirms the molecular mass of the complete conjugated peptide.

Peptide mapping: Enzymatic or chemical fragmentation followed by MS analysis helps verify the amino acid sequence and modification site.

Tandem mass spectrometry: MS/MS can distinguish sequence variants and provide evidence for attachment of the side chain at Lys26.

Chromatographic purity: RP-HPLC or UHPLC detects deletion sequences, incomplete conjugates and related substances.

Amino acid and chiral analysis: These methods support composition and stereochemical identity, particularly when non-natural residues are present.

Biological activity testing: A GLP-1 receptor-based assay evaluates functional potency but should be used together with chemical characterization.

Reference standards, system suitability criteria and method validation should match the intended research or quality-control application.

Selecting Semaglutide for Research

Before purchasing semaglutide or a related intermediate, researchers should confirm:

  • Whether the product is complete semaglutide or only the peptide backbone
  • The full annotated sequence and modification site
  • The identity of the linker and fatty diacid
  • Molecular form, counterion and water content
  • Peptide purity and net peptide content
  • Analytical methods used for identity and purity
  • Availability of HPLC and mass-spectrometry data
  • Storage, reconstitution and stability information
  • Research-grade or GMP manufacturing status
  • Intended-use restrictions

A high HPLC area percentage does not necessarily represent the actual amount of semaglutide in a sample. Water, counterions, residual solvents and non-peptide components may affect net content.

Frequently Asked Questions

Is semaglutide a peptide or a protein?

Semaglutide is generally classified as a peptide or lipopeptide. Its backbone contains 31 amino acid residues, which is substantially shorter than most proteins.

Is semaglutide identical to human GLP-1?

No. It is an analogue of GLP-1(7-37) containing Aib8 and Arg34 substitutions plus a Lys26-linked fatty diacid side chain.

How many amino acids are in semaglutide?

The peptide backbone contains 31 amino acid residues. The linker and fatty diacid are additional structural components and are not counted as amino acids in the main peptide sequence.

What does Aib mean in the sequence?

Aib means 2-aminoisobutyric acid. It is a non-natural amino acid used at position 8 to improve resistance to DPP-4 degradation.

Can semaglutide be written as a one-letter sequence?

Only partially. Aib has no standard one-letter amino acid code, and ordinary sequence notation cannot fully represent the branched Lys26 side chain. An annotated three-letter sequence or complete chemical structure is more accurate.

Is the semaglutide main chain the same as semaglutide?

No. The main chain does not include the complete Lys26 linker and C18 fatty diacid. It should be treated as a synthetic intermediate unless the full structural modification has been confirmed.

Conclusion

The semaglutide peptide sequence contains 31 amino acid residues and is closely related to human GLP-1(7-37). Its functional design depends on three defining features: Aib at position 8, arginine at position 34 and a C18 fatty diacid attached to Lys26 through a hydrophilic linker.

For research, synthesis and quality control, the peptide sequence should always be evaluated together with the side-chain structure, attachment site, stereochemistry, molecular form and analytical data. This complete structural view is essential for distinguishing authentic semaglutide from an unmodified backbone, a sequence variant or a related impurity.

References

  1. PubChem: Semaglutide
  2. DailyMed: Ozempic Prescribing Information
  3. EMA: Ozempic Assessment Report
  4. Lau J. et al. Discovery of Semaglutide

Hot Articles

Categories