2026-08-31 Posted by TideChem view:205
Semaglutide is a modified glucagon-like peptide-1 receptor agonist derived from human GLP-1(7-37). Its peptide backbone contains 31 amino acid residues, but the complete molecule cannot be described by the amino acid sequence alone. It also includes a non-natural amino acid and a fatty diacid side chain attached through a defined linker.
These structural modifications improve resistance to enzymatic degradation and promote albumin binding while retaining GLP-1 receptor activity. Understanding the complete semaglutide peptide sequence is therefore important for peptide synthesis, analytical method development, impurity identification, and pharmaceutical research.
The semaglutide peptide backbone is:
His-Aib-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys-Glu-Phe-Ile-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly-OH
The lysine residue corresponding to Lys26 in native GLP-1 carries a side chain composed of:
The spacer units may also be described as ADO, AEEA or OEG units, depending on the naming convention used by the supplier or analytical laboratory.
Semaglutide has the molecular formula C187H291N45O59 and a molecular weight of 4113.58 g/mol according to official product information. PubChem classifies it as both a polypeptide and a lipopeptide. PubChem, DailyMed
Semaglutide is based on GLP-1(7-37), so scientific literature normally uses the residue numbering of native GLP-1 rather than numbering the semaglutide chain from 1 to 31.
| Structural feature | GLP-1 numbering | Position in the 31-residue chain |
| N-terminal histidine | His7 | 1 |
| Aib substitution | Aib8 | 2 |
| Side-chain attachment | Lys26 | 20 |
| Arginine substitution | Arg34 | 28 |
| C-terminal glycine | Gly37 | 31 |
This distinction explains why Aib8 is sometimes described as the second residue of semaglutide. Both statements are correct, but they use different numbering systems.
The human GLP-1(7-37) sequence is:
His-Ala-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys-Glu-Phe-Ile-Ala-Trp-Leu-Val-Lys-Gly-Arg-Gly-OH
Semaglutide retains 29 of these 31 amino acid positions, giving it approximately 94% sequence identity with native GLP-1. It contains two amino acid substitutions and one major side-chain modification:
The European Medicines Agency describes semaglutide as an Aib8, Arg34-GLP-1(7-37) analogue with a side chain attached to Lys26. EMA assessment report
Aib is 2-aminoisobutyric acid, a non-proteinogenic amino acid. It replaces alanine at position 8, which is close to the N-terminus and involved in recognition by dipeptidyl peptidase-4.
This substitution increases resistance to DPP-4-mediated degradation. Aib does not have a standard one-letter amino acid code, so a conventional one-letter sequence cannot represent semaglutide accurately without additional annotation.
Native GLP-1 contains lysine residues at positions 26 and 34. Replacing Lys34 with arginine leaves Lys26 as the principal site for controlled side-chain attachment.
This modification helps prevent the formation of differently acylated positional variants during manufacturing.
The ε-amino group of Lys26 is connected to a C18 fatty diacid through a hydrophilic spacer. The spacer helps separate the lipid group from the receptor-binding peptide region, while the fatty diacid supports reversible albumin association.
Albumin binding reduces rapid renal clearance and contributes to the prolonged pharmacokinetic profile of semaglutide. The importance of the fatty acid and linker design was described in the original semaglutide discovery research. Journal of Medicinal Chemistry
A sequence containing 31 amino acids is not necessarily complete semaglutide. The identity of the molecule also depends on:
An unmodified semaglutide backbone, sometimes described as a semaglutide main chain or des-acyl semaglutide, is an intermediate rather than the complete active molecule.
Likewise, an amino acid sequence, a research-grade peptide and an approved pharmaceutical product should not be treated as equivalent. Pharmaceutical equivalence also requires appropriate manufacturing controls, formulation, stability data, biological testing and regulatory authorization.
Regulatory documents describe commercial semaglutide production as yeast fermentation followed by chemical modification and purification. Research-scale material may also be prepared using solid-phase peptide synthesis, fragment condensation or hybrid approaches.
Important manufacturing challenges include:
The lipid side chain can alter chromatographic behavior and solubility. Purification methods developed for ordinary hydrophilic peptides may therefore require different gradients, stationary phases or sample-preparation conditions.
No single analytical result is sufficient to establish the complete identity of semaglutide. A suitable characterization package may include:
Intact-mass analysis: LC-MS or high-resolution MS confirms the molecular mass of the complete conjugated peptide.
Peptide mapping: Enzymatic or chemical fragmentation followed by MS analysis helps verify the amino acid sequence and modification site.
Tandem mass spectrometry: MS/MS can distinguish sequence variants and provide evidence for attachment of the side chain at Lys26.
Chromatographic purity: RP-HPLC or UHPLC detects deletion sequences, incomplete conjugates and related substances.
Amino acid and chiral analysis: These methods support composition and stereochemical identity, particularly when non-natural residues are present.
Biological activity testing: A GLP-1 receptor-based assay evaluates functional potency but should be used together with chemical characterization.
Reference standards, system suitability criteria and method validation should match the intended research or quality-control application.
Before purchasing semaglutide or a related intermediate, researchers should confirm:
A high HPLC area percentage does not necessarily represent the actual amount of semaglutide in a sample. Water, counterions, residual solvents and non-peptide components may affect net content.
Semaglutide is generally classified as a peptide or lipopeptide. Its backbone contains 31 amino acid residues, which is substantially shorter than most proteins.
No. It is an analogue of GLP-1(7-37) containing Aib8 and Arg34 substitutions plus a Lys26-linked fatty diacid side chain.
The peptide backbone contains 31 amino acid residues. The linker and fatty diacid are additional structural components and are not counted as amino acids in the main peptide sequence.
Aib means 2-aminoisobutyric acid. It is a non-natural amino acid used at position 8 to improve resistance to DPP-4 degradation.
Only partially. Aib has no standard one-letter amino acid code, and ordinary sequence notation cannot fully represent the branched Lys26 side chain. An annotated three-letter sequence or complete chemical structure is more accurate.
No. The main chain does not include the complete Lys26 linker and C18 fatty diacid. It should be treated as a synthetic intermediate unless the full structural modification has been confirmed.
The semaglutide peptide sequence contains 31 amino acid residues and is closely related to human GLP-1(7-37). Its functional design depends on three defining features: Aib at position 8, arginine at position 34 and a C18 fatty diacid attached to Lys26 through a hydrophilic linker.
For research, synthesis and quality control, the peptide sequence should always be evaluated together with the side-chain structure, attachment site, stereochemistry, molecular form and analytical data. This complete structural view is essential for distinguishing authentic semaglutide from an unmodified backbone, a sequence variant or a related impurity.