How to read a peptide sequence
The one- and three-letter codes, which end is which, what Ac- and -NH2 mean, how D-residues, rings and lipids are written, and four worked examples from the catalog.
A peptide sequence is a list of amino-acid residues written from the N-terminus on the left to the C-terminus on the right, in either a one-letter code (GEPPPGKPADDAGLV) or a three-letter code (Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val). The rules come from the IUPAC-IUB Joint Commission on Biochemical Nomenclature, published in 1983 and still in force. With the codes, the two ends and a handful of modification symbols, any product page in the catalog can be read.
The alphabet
Twenty amino acids are encoded by the genetic code, and each has a three-letter symbol and a one-letter symbol. The three-letter symbol is usually the first three letters of the name; the one-letter symbols were assigned by IUPAC-IUB using initial letters where there was no clash and arbitrary letters where there was.
| One | Three | Amino acid | One | Three | Amino acid |
|---|---|---|---|---|---|
| A | Ala | alanine | M | Met | methionine |
| C | Cys | cysteine | N | Asn | asparagine |
| D | Asp | aspartic acid | P | Pro | proline |
| E | Glu | glutamic acid | Q | Gln | glutamine |
| F | Phe | phenylalanine | R | Arg | arginine |
| G | Gly | glycine | S | Ser | serine |
| H | His | histidine | T | Thr | threonine |
| I | Ile | isoleucine | V | Val | valine |
| K | Lys | lysine | W | Trp | tryptophan |
| L | Leu | leucine | Y | Tyr | tyrosine |
Three more letters appear in older sequence data: B for "Asp or Asn," Z for "Glu or Gln" when the two could not be distinguished, and X for an unknown or unusual residue. A synthetic peptide sold with a certificate should never need them.
Which end is which
The IUPAC-IUB rule for the one-letter system is explicit: "The letter written at the left-hand end is that of the amino-acid residue carrying the free amino group, and the letter written at the right-hand end is that of the residue carrying the free carboxyl group." The residue with the free amino group is the N-terminal residue; the one with the free carboxyl group is C-terminal. Three-letter formulas follow the same order, "with the N-terminal residue on the left, and the C-terminal on the right."
Every hyphen in a three-letter formula is a peptide bond: Gly-Gly-Gly means three glycines joined by two amide bonds. A symbol with hyphens on both sides (-Gly-) is an internal residue; a symbol with no hyphen on the left has a free amine, and one with no hyphen on the right has a free carboxyl. When writers want to make the free ends explicit they write H-Met-Glu-His-Phe-Pro-Gly-Pro-OH, where H- stands for the hydrogen on the N-terminal amine and -OH for the hydroxyl of the C-terminal carboxylic acid. PubChem lists Semax in exactly that form.
Same letters, different molecule
Order is identity. IUPAC-IUB's own example is the pair glycylalanine (Gly-Ala) and alanylglycine (Ala-Gly): the same two amino acids, condensed in opposite directions, giving two different compounds. Both have the same formula and the same mass, so a mass spectrometer cannot tell them apart. In general, mass spectrometry confirms composition, and only a sequencing method or a chromatographic comparison against a reference standard confirms order. Mass spectrometry identity covers what a mass match proves, and Reading a chromatogram what the HPLC trace can and cannot show.
A sequence printed on a label is therefore a claim about order, and the certificate's identity test is what backs it up.
The ends: Ac- and -NH2
The two most common modifications on product pages sit at the ends of the chain.
Ac- at the left means the N-terminal amine has been acetylated. IUPAC-IUB lists Ac- as the symbol for the acetyl substituent. An acetylated N-terminus has no free amine, so the residue is still called N-terminal (the rules allow it to be "acetylated or formylated") but the chain no longer carries a positive charge at that end.
-NH2 at the right means the C-terminal carboxyl has been converted to a primary amide. IUPAC-IUB describes the C-terminal residue as one that "may, for example, acylate ammonia to give -NH-CHR-CO-NH2." Their example is thyroliberin, Glp-His-Pro-NH2, where Glp is pyroglutamate, itself a cyclised glutamine. An amidated C-terminus has no free carboxylic acid and therefore no negative charge at that end.
Both changes alter the mass (acetylation adds 42 daltons, amidation replaces OH with NH2 and subtracts about one dalton) and the charge, and both are made deliberately during synthesis. An analogue with Ac- and -NH2 is a different molecule from the free acid with the same letters, and the certificate's mass should reflect that.
Melanotan I (afamelanotide) shows both at once. PubChem gives its sequence as Ac-Ser-Tyr-Ser-Nle-Glu-His-D-Phe-Arg-Trp-Gly-Lys-Pro-Val-NH2: thirteen residues, acetylated at the start, amidated at the end, and carrying two further modifications described next. It is in the catalog as Melanotan I.
D-amino acids and unusual residues
The rules state that "residue symbols written in a sequence denote the L configuration for chiral amino acids, unless otherwise indicated." A D residue "is shown by inserting a D before the symbol, separated from it by a hyphen," as in D-Phe. In print the D is a small capital; on a web page it is usually just a capital D. The one-letter system has no symbol for configuration, which is why peptides containing D-residues are almost always written in three-letter form.
Swapping an L-residue for its D mirror image changes the shape of the chain at that point without changing the mass, so a D-substitution is invisible to mass spectrometry. Afamelanotide's D-Phe at position 7 is one example.
Non-standard residues get their own symbols, which the rules say "should be defined in each publication." Nle in afamelanotide is norleucine, a straight-chain isomer of leucine. Glp is pyroglutamate. Aib, Orn and dozens of others appear in the wider literature. When a product page uses one, the symbol should be spelled out somewhere on the page.
Rings, disulfides and lipids
Cyclic peptides. A ring closed by a normal peptide bond is written with the sequence in parentheses preceded by cyclo, in IUPAC-IUB's example cyclo(-Val-Orn-Leu-D-Phe-Pro-Val-Orn-Leu-D-Phe-Pro-) for gramicidin S. A ring closed through a side chain is a lactam bridge. The Vyleesi prescribing information writes bremelanotide as Ac-Nle-cyclo-(Asp-His-D-Phe-Arg-Trp-Lys-OH), meaning the side-chain carboxyl of Asp and the side-chain amine of Lys are joined, closing a ring that leaves the Nle at the start and the Lys carboxyl at the end outside it.
Disulfides. Two cysteines can be joined through their sulfur atoms. In systematic names this appears as "cyclic (2→7)-disulfide" or similar, giving the positions of the two Cys residues. Amylin and its analogue cagrilintide carry a Cys2 to Cys7 disulfide.
Side-chain links. Glutathione, gamma-Glu-Cys-Gly, is IUPAC-IUB's example of a bond made through a side-chain carboxyl rather than the backbone; the Greek gamma tells you which carboxyl.
Lipids and polymers. A fatty acid or a polyethylene glycol chain is written as a substituent on the residue that carries it, usually a lysine side chain or the N-terminus, and the linker between them is named too. Semaglutide's fatty diacid on a lysine is described in Thirty-nine residues and a fatty acid. These additions are the largest single source of mass difference between a parent hormone and its analogue.
Four worked examples from the catalog
BPC-157. Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val, or GEPPPGKPADDAGLV. Fifteen residues, so a pentadecapeptide. Free N-terminus and free C-terminus. Three prolines in a row near the start, and an Asp-Asp pair in the middle. Molecular weight 1419.5, CAS 137525-51-0, PubChem CID 9941957. Stocked as BPC-157, a lyophilized vial.
Semax. Met-Glu-His-Phe-Pro-Gly-Pro, or MEHFPGP. Seven residues, a heptapeptide. PubChem's title for it, "ACTH (4-7), Pro-Gly-Pro-," tells you its design: residues 4 to 7 of adrenocorticotropic hormone with a Pro-Gly-Pro tail added. The N-terminal methionine is the residue to watch for oxidation. Molecular weight 813.9, CAS 80714-61-0, CID 9811102. Stocked as Semax.
MOTS-c. Met-Arg-Trp-Gln-Glu-Met-Gly-Tyr-Ile-Phe-Tyr-Pro-Arg-Lys-Leu-Arg, or MRWQEMGYIFYPRKLR. Sixteen residues, with two methionines and a tryptophan, all oxidisable, and three arginines that make it strongly basic. PubChem lists it as H-MRWQEMGYIFYPRKLR-OH, confirming free ends. Molecular weight 2174.6, CAS 1627580-64-6, CID 146675088. Stocked as MOTS-c.
Epitalon. Ala-Glu-Asp-Gly, or AEDG. Four residues, a tetrapeptide, two of them acidic. Molecular weight 390.35, CAS 307297-39-8, CID 219042. Stocked as Epitalon; its naming history is in Epithalon, epitalon and the word bioregulator.
What a sequence does not tell you
A sequence is the identity of the molecule and nothing more. It does not say how pure the material in the vial is, what fraction of the mass is peptide rather than water and counter-ion, or which salt form was supplied. Those are certificate questions, covered in Net peptide content and What a peptide is. It also does not say anything about what the molecule does; the site does not publish that.
Sources
- Nomenclature and Symbolism for Amino Acids and Peptides, IUPAC-IUB Joint Commission on Biochemical Nomenclature, Recommendations 1983, web edition
- Sections 3AA-11 to 3AA-13, Peptide Nomenclature, IUPAC-IUB JCBN, 1983
- Sections 3AA-14 to 3AA-16, The Three-Letter System, IUPAC-IUB JCBN, 1983
- Sections 3AA-18 and 3AA-19, Symbols for Substituents and Peptide Symbolism, IUPAC-IUB JCBN, 1983
- Sections 3AA-20 and 3AA-21, The One-Letter System, IUPAC-IUB JCBN, 1983
- BPC-157, PubChem CID 9941957, National Library of Medicine, accessed 2026-09-21
- Semax, PubChem CID 9811102, National Library of Medicine, accessed 2026-09-21
- MOTS-c, PubChem CID 146675088, National Library of Medicine, accessed 2026-09-21
- Epitalon, PubChem CID 219042, National Library of Medicine, accessed 2026-09-21
- Afamelanotide, PubChem CID 16197727, National Library of Medicine, accessed 2026-09-21
- Bremelanotide, PubChem CID 9941379, National Library of Medicine, accessed 2026-09-21
- Vyleesi (bremelanotide) prescribing information, DailyMed, National Library of Medicine, accessed 2026-09-21
For laboratory research use only. Not a drug, not a supplement, and nothing here is a claim about what any of this material does in a person or an animal.

