A peptide sequence is written from the N-terminus on the left to the C-terminus on the right, using either one-letter or three-letter amino-acid codes, with any chemical modification written as a prefix or suffix at the end it applies to. Learning how to read a peptide sequence takes about ten minutes and it is the single fastest way to sanity-check a vial against its certificate of analysis: the sequence encodes the identity, it predicts the molecular weight, and it explains most of what a molecule does in solution. This guide covers the two code systems, direction conventions, the modification shorthand you will actually meet on research labels, and a worked calculation that turns a sequence into a molecular weight you can compare against a COA. Everything below describes laboratory reference material supplied for research use only.
Direction: why N-to-C matters
Amino acids join through a peptide bond formed between the carboxyl group of one residue and the amino group of the next. That reaction is directional, so a chain has two chemically distinct ends: a free amino group (the N-terminus) and a free carboxyl group (the C-terminus). By universal convention the N-terminus is written first. Gly-Glu-Pro and Pro-Glu-Gly are different molecules with identical composition and identical mass, and a mass spectrometer alone will not separate them, which is why MS identity confirmation is normally paired with a stated sequence rather than used instead of one.
Residue numbering follows the same direction. When a label reads HGH Fragment 176-191, the numbers refer to positions counted from the N-terminus of the parent 191-residue protein; the fragment is the 16 residues occupying those positions, not a separately designed sequence.
How to read a peptide sequence: one-letter and three-letter codes
Three-letter codes are readable and unambiguous, so they dominate product labels and COAs. One-letter codes are compact and dominate databases, papers and synthesis order forms. Both describe the same twenty proteinogenic residues.
| Amino acid | 3-letter | 1-letter | Average residue mass (Da) |
|---|---|---|---|
| Glycine | Gly | G | 57.05 |
| Alanine | Ala | A | 71.08 |
| Serine | Ser | S | 87.08 |
| Proline | Pro | P | 97.12 |
| Valine | Val | V | 99.13 |
| Threonine | Thr | T | 101.10 |
| Cysteine | Cys | C | 103.14 |
| Leucine / Isoleucine | Leu / Ile | L / I | 113.16 |
| Asparagine | Asn | N | 114.10 |
| Aspartic acid | Asp | D | 115.09 |
| Glutamine | Gln | Q | 128.13 |
| Lysine | Lys | K | 128.17 |
| Glutamic acid | Glu | E | 129.12 |
| Methionine | Met | M | 131.19 |
| Histidine | His | H | 137.14 |
| Phenylalanine | Phe | F | 147.18 |
| Arginine | Arg | R | 156.19 |
| Tyrosine | Tyr | Y | 163.18 |
| Tryptophan | Trp | W | 186.21 |
Leucine and isoleucine share a mass because they are structural isomers. That single fact explains a recurring frustration in peptide analytics: standard mass spectrometry cannot distinguish L from I, so their assignment rests on the synthesis record and on chromatographic behaviour rather than on mass alone.
Reading a real label
BPC-157 is catalogued as Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val, which in one-letter code is GEPPPGKPADDAGLV. Fifteen residues, hence "pentadecapeptide". Three consecutive prolines near the N-terminus give the chain a rigid kink, and the two adjacent aspartates carry negative charge at neutral pH — both structural facts you can read straight off the sequence, and both relevant to solubility behaviour.
Modification shorthand
Most research peptides are not plain unmodified chains. The notation is consistent once you know it.
- Ac- at the left end means N-terminal acetylation: the free amino group is capped with an acetyl group, adding 42.04 Da and removing a positive charge. TB-500 is written Ac-Leu-Lys-Lys-Thr-Glu-Thr-Gln, i.e. Ac-LKKTETQ.
- -NH2 at the right end means C-terminal amidation: the terminal -OH is replaced by -NH2, subtracting 0.98 Da and removing a negative charge. Many natural signalling peptides are amidated in vivo, and the amide is usually required for receptor activity — oxytocin is written Cys-Tyr-Ile-Gln-Asn-Cys-Pro-Leu-Gly-NH2.
- D- before a residue marks the D-enantiomer, as in the D-Trp of GHRP-6 (His-D-Trp-Ala-Trp-D-Phe-Lys-NH2). D-residues have the same mass as their L-counterparts but are poor substrates for mammalian proteases, so they are inserted deliberately to slow degradation.
- Non-standard residues appear by name. Aib is α-aminoisobutyric acid, a helix-stabilising, protease-resistant residue used in Ipamorelin (Aib-His-D-2-Nal-D-Phe-Lys-NH2) and at position 2 of the GLP-1 backbone in semaglutide and tirzepatide. 2-Nal is 2-naphthylalanine.
- Brackets and superscripts denote substitutions relative to a parent sequence, e.g. [D-Ala²] means position 2 has been swapped for D-alanine.
- Cyclisation is shown by a line, a "cyclo(...)" wrapper, or a stated Cys–Cys disulfide. Oxytocin's Cys1 and Cys6 form a disulfide bridge that closes a six-residue ring; that bond is reduction-sensitive and is a real storage consideration.
Worked example: sequence to molecular weight
Molecular weight is the sum of the residue masses plus 18.02 Da for the water molecule released on chain formation, then adjusted for modifications. Two examples using catalogued values:
- KPV (Lys-Pro-Val). 128.17 + 97.12 + 99.13 = 324.42. Add 18.02 → 342.44 Da. The catalogue lists 342.43; the 0.01 difference is rounding in the residue table.
- Epitalon (Ala-Glu-Asp-Gly, AEDG). 71.08 + 129.12 + 115.09 + 57.05 = 372.34. Add 18.02 → 390.36 Da, against a catalogued 390.35.
- Add a modification. An acetylated, amidated version of the same tetrapeptide would be 390.36 + 42.04 − 0.98 = 431.42 Da.
Agreement within about 0.1% is the expected result. A discrepancy of 18 Da usually means someone forgot the water term; 42 Da means an acetyl group was ignored; 1 Da means an amide was missed. Once you have a reliable molecular weight you can convert between mass and moles for buffer preparation — the arithmetic is covered in molecular weight, moles and molarity and in the molarity calculator.
Sequences on a COA versus sequences in a catalogue
A catalogue entry states the intended sequence. A COA reports what was measured: an observed monoisotopic or average mass from MS, and a purity figure from reversed-phase HPLC. Neither technique reads the sequence residue by residue unless MS/MS fragmentation data are supplied, which is uncommon on routine commercial certificates. Treat the stated sequence as the manufacturer's declaration and the measured mass as the check on it.
Common misreadings
- Reading C-to-N. Reversing a sequence produces a different molecule (a retro-peptide) with the same mass. Always confirm the leftmost residue is the N-terminus.
- Treating a salt form as part of the sequence. Trifluoroacetate or acetate counter-ions are not residues; they add mass to the vial contents but not to the peptide. This is the usual reason a weighed amount of powder contains less peptide than the label mass implies, and it is why TFA salt content belongs on the COA.
- Assuming fragment numbering matches the fragment length. "176-191" is 16 residues, not 191.
- Ignoring the amide. A sequence that is active only in its amidated form is a different research reagent from the free-acid version, even though the two differ by one dalton.
- Confusing blends with single sequences. A co-formulated vial such as the KLOW Blend has four sequences and no single molecular weight; see blends vs single vials.
When a sequence, a stated molecular weight and an observed MS mass all agree, you have reasonable evidence that the vial contains what the label says. When they disagree, the sequence is almost always the fastest place to find out why.