A peptide is a molecule built from two or more amino acid units joined through amide bonds, with the carbonyl carbon of one unit bonded to the nitrogen of the next. That single sentence answers the definition question, but it leaves three practical questions open. A chain drawing shows which atoms are connected. A sequence string shows the order of the units and the direction of the chain. A material label on a research vial tells you what a supplier calls an item, which may or may not carry enough detail to reconstruct either of the first two. This page teaches the first two, so that a reader can decode a structure or a sequence and then judge what a label leaves out.
The scope is deliberately narrow. Everything below is about composition, connectivity, and notation. Nothing here describes how a peptide is made, what it does in a living system, or how to judge a certificate. Those are separate documentation tasks with separate owners on this site.
What makes a molecule a peptide?
The defining feature is the bond, not the size. In the nomenclature recommendations maintained by the IUPAC-IUB Joint Commission on Biochemical Nomenclature, a peptide is an amide formed from two or more amino carboxylic acid molecules, connected by a covalent bond from the carbonyl carbon of one to the nitrogen atom of another. Chemists call that specific C(=O)-N connection a peptide bond. The molecules on either side can be identical or different, and the count can be two or two hundred.
A peptide is a molecule in which two or more amino acid residues are connected by amide bonds, each formed between the carboxyl carbon of one residue and the amino nitrogen of the next. The most familiar peptides are built from alpha-amino acids joined through their alpha groups, but the definition covers amide connections between amino carboxylic acids generally, so the alpha pattern is the common case rather than the rule.
Two consequences follow. First, there is no minimum or maximum length written into the definition. A two-residue molecule is a peptide, and so is a chain long enough that most chemists would casually call it a protein; the size words are conventions layered on top, and they get their own section below. Second, the definition does not require that the amide bond involve the alpha carboxyl group and the alpha amino group. When the linkage runs through a side-chain group instead, the molecule is still a peptide, and the name has to say so. Glutathione is the standard teaching case, and this page returns to it.
What the definition excludes is just as useful. A mixture of free amino acids in a container is not a peptide, because nothing is bonded. A list of amino acid names on a label is not a structure, because a list carries no connectivity. The word peptide refers to a covalent arrangement, so the test for whether a drawing shows a peptide is whether you can trace a C(=O)-N bond between two residues.
How do amino acids become residues in a peptide chain?
To read a peptide structure, start with one free alpha-amino acid. The IUPAC recommendations draw it around a single carbon, the alpha carbon, which carries four groups: an amino group, a carboxyl group, a hydrogen atom, and a side chain that varies from one amino acid to the next. That layout is what the phrase alpha-amino acid means: the amino group sits on the carbon adjacent to the carboxyl carbon.
side chain (R)
|
H2N ---- C(alpha) ---- COOH one free alpha-amino acid
|
HWhen two amino acids are joined into a peptide, the atoms of one water molecule are formally absent from the product: an OH from one carboxyl group and an H from the other amino group are not in the chain. This is chemical bookkeeping, and it matters for reading formulas. Because those atoms are gone, what remains of each amino acid inside the chain is not the original molecule. The recommendations give it a separate name: an amino-acid residue. A residue is the part of an amino acid that persists in the assembled peptide.
The distinction sounds pedantic until you compare formulas. The molecular formula of a free amino acid and the contribution of that same amino acid as a residue differ by the atoms of water. If a document lists the formula of a dipeptide, that formula is smaller than the sum of the two free amino acids. A reader who expects the formulas to add up will conclude that something is missing when nothing is.
Once residues are connected, a repeating pattern appears along the chain. Each residue contributes three backbone atoms in order: the amino nitrogen, the alpha carbon, and the carbonyl carbon. The recommendations call this repeating N, C-alpha, C(=O) sequence the main chain. Everything else attached to each alpha carbon is a side chain. Backbone atoms are the same for every residue; side chains are what make glycine different from lysine.
One caution about the drawings. The nomenclature tables present amino acids as neutral structural formulas because that is the convention for naming. The recommendations are explicit that this is a drawing convention. It is not a claim that the neutral form is the predominant species in water, in a solid, or in any research material. When you copy a structure from a naming source, you inherit that convention, and a documentation reader should not read charge information into it.
How do you find the peptide bond in a structural formula?
In a drawn structure, the peptide bond is the single bond between a carbonyl carbon, drawn as C with a double-bonded O, and the nitrogen of the next residue. To find it, locate every C(=O)-N pair along the backbone. A dipeptide has one such pair; a tripeptide has two. Counting these bonds is the fastest way to count residues from a drawing, because in an unbranched linear chain the number of residues is always one more than the number of peptide bonds.
The bond and the residue are different things, and mixing them up produces wrong descriptions. The peptide bond is a few atoms wide: carbonyl carbon, carbonyl oxygen, nitrogen, and the hydrogen on that nitrogen. The residue is the entire unit from one backbone nitrogen through its carbonyl carbon, including its side chain. When a document says that a molecule contains a glycine residue, it is describing a unit. When it says that two residues are linked by a peptide bond, it is describing a connection.
Here is a worked example built only from the two smallest common amino acids. Glycine has a hydrogen atom in place of a carbon-containing side chain. Alanine has a methyl group. Join them in one order and you get a dipeptide whose N-terminal residue is glycine and whose C-terminal residue is alanine. Join them in the other order and both residues swap positions. The recommendations name these glycylalanine and alanylglycine, and they are different compounds.
Gly-Ala (glycylalanine)
H2N - CH2 - C(=O) - NH - CH(CH3) - COOH
[Gly residue] ^ [Ala residue]
|
peptide bond
Ala-Gly (alanylglycine)
H2N - CH(CH3) - C(=O) - NH - CH2 - COOH
[Ala residue] ^ [Gly residue]
|
peptide bondNotice what the drawing gives you that a list would not. A label reading glycine and alanine tells you two components. It does not tell you whether they are bonded, and if bonded, in which order. The drawing settles both. This is why a structural formula is the highest-resolution description a document can carry, and why a sequence string is the next best thing.
What do N-terminus and C-terminus mean in a peptide sequence?
A linear peptide has two ends that are chemically different. One end carries the residue whose amino nitrogen is not part of a peptide bond; this is the N-terminal residue. The other end carries the residue whose carboxyl carbon is not part of a peptide bond; this is the C-terminal residue. Every residue in between is an internal residue, bonded on both sides.
The naming convention fixes a reading direction. Names and sequences run from the N-terminal residue on the left to the C-terminal residue on the right. In a systematic name, every residue except the last is written with an -yl ending, marking it as an acyl unit bonded to the next residue, and the C-terminal residue keeps its full amino-acid name. That is why glycylalanine ends in alanine: alanine is the C-terminal residue.
Peptide sequences are written from the N-terminus on the left to the C-terminus on the right. Reversing a sequence is not a cosmetic change. It moves a different residue to each end and changes which residue's carbonyl carbon is bonded to which residue's nitrogen. Gly-Ala and Ala-Gly contain the same two residues and one peptide bond each, yet they are different compounds because their connectivity differs.
N-terminus C-terminus
(amino end) (carboxyl end)
| |
v v
[Res 1] --- [Res 2] --- [Res 3] --- ... --- [Res n]
first internal residues last
position numbers count from the N-terminal residueModified ends do not remove the ends. A peptide can carry an acyl group on its N-terminal nitrogen or an amide in place of its C-terminal carboxyl group. The recommendations keep the terms N-terminal and C-terminal for those residues regardless; the terms describe position in the chain, not whether the terminal group is unmodified. A documentation reader should therefore expect a modified terminus to be written as an annotation on the terminal residue, not as an absence of a terminus.
How do peptide names, three-letter codes, and one-letter sequences differ?
Three systems describe the same chain at three levels of compression. A systematic name spells out each residue in order. A three-letter sequence uses a symbol per residue. A one-letter sequence uses a single character per residue. Each was designed for a different job, and each drops information that the others keep.
Three-letter symbols follow a fixed pattern: one uppercase letter followed by two lowercase letters. Gly, Ala, Cys, Glu, and Lys are the symbols for the five amino acids used as examples on this page. In a peptide representation, hyphens carry meaning. A symbol with a hyphen on one side or both sides denotes a residue bonded at those positions; a symbol with no hyphens denotes the free amino acid. So Gly-Gly-Gly describes one tripeptide containing three glycine residues, not three separate glycine molecules. The recommendations also state that a standard symbol for a chiral amino acid denotes the L configuration unless the writer adds a contrary annotation, and that any symbol outside the standard set must be defined in the document that uses it.
| Name | Three-letter symbol | One-letter symbol | Side-chain feature as described in the source |
|---|---|---|---|
| Glycine | Gly | G | No carbon-containing side chain; a hydrogen occupies the side-chain position |
| Alanine | Ala | A | Methyl group |
| Cysteine | Cys | C | Side chain contains a sulfur atom |
| Glutamic acid | Glu | E | Side chain carries an additional carboxyl group |
| Lysine | Lys | K | Side chain carries an additional amino group |
One-letter symbols exist because long sequences are unreadable in any other form. The recommendations describe the compact notation as intended for long sequences, for aligning sequences against each other, and for computer representation. Its cost is transparency. A reader who does not know the code cannot recover the amino acid from the letter, so explanatory text should define the symbols it uses. The code also contains placeholders. B stands for an aspartic acid or asparagine residue whose identity has not been resolved; Z plays the same role for glutamic acid and glutamine; X marks an unknown or other residue. A fragment notation exists to signal that a shown end has not been demonstrated to be the actual terminus of the molecule.
Two confusions recur in documentation. The first is between a residue number and an atom locant. Residue 3 means the third residue counting from the N-terminus; it says nothing about a specific atom. A locant such as the epsilon nitrogen of a lysine residue names an atom within a residue. The second is between peptide one-letter codes and nucleotide codes. The letters G, A, and C appear in both systems and mean different things. A string of letters is only a peptide sequence when the document says it is.
What does each representation of one peptide show and omit?
The recommendations note that structural formulas and symbols can be combined when abbreviated notation would hide important detail. The table below is an editorial comparison built from the definitions above. It asks the same three questions of each representation: what does it show, what does it omit, and what must a reader check elsewhere.
| Representation | Information visible | Information absent | What a reader must check |
|---|---|---|---|
| Structural formula | Every atom and bond, including side chains, termini, and non-alpha linkages | Nothing about the material; stereochemistry only if wedges or descriptors are drawn | Whether stereodescriptors are present and whether the drawing follows the neutral naming convention |
| Three-letter sequence | Residue identity and order; hyphens mark bonds; annotations can mark modified termini and side-chain links | Atom-level detail; configuration is assumed L unless annotated | Any nonstandard symbol and any missing annotation on a terminus or linkage |
| One-letter sequence | Residue order in the most compact form; suited to alignment and software | Modifications, linkage type, configuration, and terminal state unless a separate annotation layer is supplied | Whether placeholders such as B, Z, or X appear and whether a fragment marker is present |
| Molecular formula | Atom counts by element | Order, connectivity, and configuration; two different sequences can share one formula | Whether the formula describes the free peptide or a salt or other associated form |
| Short trivial name | A memorable label that points to one specific structure by convention | Everything structural; the name is only useful if the reader already knows which structure it points to | The defining reference that fixes the structure behind the name |
Read across a row and the pattern is consistent: each step toward compactness removes a class of information. Read down the last column and you get the questions to ask of any document that supplies only one representation. A document that gives a trivial name and nothing else has given you a pointer, not a description, and the worksheet later on this page is built to record that gap rather than fill it from memory.
What is the difference between a dipeptide, tripeptide, polypeptide, and protein?
Size words in peptide chemistry are counts of residues. A dipeptide has two residues, a tripeptide three, a tetrapeptide four, and so on through the numerical prefixes. The count is of residues, not of letters in a name, hyphens in a symbol string, or amino acids in an inventory. Gly-Gly-Gly is a tripeptide because it contains three residues; its trivial name, triglycine, happens to make the count obvious, but names are not always that cooperative.
Names such as dipeptide, tripeptide, and tetrapeptide state the exact number of amino acid residues in a chain. The broader words oligopeptide, polypeptide, and protein describe size ranges by convention only; the IUPAC-IUB recommendations present these length terms as approximate and note that authors disagree about where the boundaries fall. No single residue count separates a peptide from a protein.
The vaguer terms are where documentation goes wrong. Oligopeptide is used for chains with a small number of residues, polypeptide for chains with many, and protein for large polypeptides, often with a functional sense attached. The recommendations present these as approximate conventions and explicitly record that authors differ on the thresholds. A reader should therefore never infer a residue count from the word polypeptide, and should never argue that an item is or is not a peptide on the basis of a length cutoff that the nomenclature body itself declines to fix.
For neutral practice, count from the structure. Given a drawing, count the C(=O)-N backbone bonds and add one. Given a three-letter string, count the symbols separated by hyphens. Given a one-letter string, count the characters, remembering that a placeholder such as X still occupies one residue position. If two documents disagree on the count, one of them is describing a different molecule, and that discrepancy is the finding to record.
How can peptide structure differ without changing the familiar short name?
A short name or a bare sequence can hide four kinds of structural difference. Each has its own notation, and a description is only unambiguous when all four are stated or explicitly marked as not applicable.
The first is configuration at the alpha carbon. Every standard amino acid except glycine has a chiral alpha carbon, and glycine is achiral because two of its four substituents are hydrogen. The D and L descriptors used in amino acid nomenclature describe configuration relative to a reference compound, as the recommendations define them; they are not statements about which direction a solution rotates polarized light. A separate sign, plus or minus, records optical rotation, and the two labels answer different questions. The R and S descriptors form a third, independent system. For most standard amino acids the L configuration corresponds to S, but the recommendations flag cysteine as the explicit exception, so a writer must not mechanically replace every L with S. Finally, the prefix DL describes an equimolar mixture of the two enantiomers in the stated context, which is a statement about a sample rather than about one molecule.
The second is the state of the termini. As described earlier, an N-terminal acyl group or a C-terminal amide leaves the residue in place but changes the molecule. A bare sequence written without terminal annotations is silent on this point, and silence is not the same as unmodified termini.
The third is the linkage position. The common peptide bond joins the alpha carboxyl group of one residue to the alpha amino group of the next. When a residue with a second carboxyl or amino group in its side chain participates through that side-chain group instead, the connectivity changes and the name must show it. The recommendations use glutathione as the example: its descriptive name is gamma-glutamylcysteinylglycine, and the gamma tells you that the glutamic acid residue is linked through its side-chain carboxyl group rather than its alpha carboxyl group. Written as a plain three-letter string without that annotation, the sequence would describe a different molecule. This site's glutathione guide covers that material's own documentation; here it serves only as the clearest example of why linkage annotation matters.
The fourth is modification of individual residues. A base string of amino acid letters does not, on its own, encode modifications to residues. The HUPO Proteomics Standards Initiative maintains ProForma, a notation standard for representing peptidoforms and proteoforms, precisely because a plain sequence cannot carry this information. Its published version table lists ProForma 2.0 as final on February 3, 2022 and 2.1 as final on June 9, 2026. This page does not teach that syntax. The point is that a sequence needs an accompanying annotation layer before it is a complete chemical description.
A one-letter or three-letter sequence identifies residue order but does not by itself fix configuration, terminal state, linkage position, or residue modification. An unambiguous chemical description therefore pairs the sequence with explicit stereodescriptors, terminal annotations, any non-alpha linkage markers, and a modification notation such as ProForma where modifications exist. A molecular formula, a sequence, and a stereochemical description each convey different information, and none substitutes for the others.
A structure-reading worksheet for research documentation
The worksheet below turns the preceding sections into a fixed set of fields. It is a reading aid for molecular descriptions in supplier documents, database records, and literature. Every field accepts the value not specified. Writing not specified is the correct output when a document is silent; inventing a plausible value is the error the worksheet exists to prevent.
| Field | What to record | Teaching example entry |
|---|---|---|
| Name as written | The exact name or code string from the document, unedited | Gly-Ala (teaching string, not a catalog item) |
| Ordered residues | Residue symbols from N-terminus to C-terminus, with the count | Gly, Ala; two residues |
| Stated termini | Whether the N-terminal and C-terminal groups are described as unmodified, modified, or not described | not specified |
| Linkage annotation | Any marker showing a non-alpha linkage, or an explicit statement that all bonds are alpha | not specified |
| Configuration information | D, L, DL, R, or S descriptors exactly as written, or none present | not specified |
| Modification notation | Any residue-level modification string and the standard it follows | not specified |
| Source identifier | The document, page, record, or accession that supplied the description | placeholder |
| Unresolved fields | Every field above that remains not specified | termini, linkage, configuration, modification |
Two habits make the worksheet reliable. First, copy strings exactly, including case and punctuation, because a lowercase letter or a missing hyphen can change meaning. Second, keep a separate line for anything you inferred rather than read. An inference that every residue is L because no annotation is present is consistent with the recommendations, but it is still an inference, and it belongs in the unresolved column until a source states it.
Nothing in the worksheet should be read as a claim about any Glow item. The example string is a teaching example chosen because it is the simplest chiral dipeptide, not because it corresponds to a product or lot.
Frequently asked questions
What is the difference between an amino acid and an amino-acid residue?
An amino acid is the complete free molecule, with its own amino group and carboxyl group intact. An amino-acid residue is what remains of that molecule after it has been incorporated into a peptide, where the atoms of water are formally absent and at least one of its groups is now part of a peptide bond. The two differ in formula, and a symbol with hyphens denotes the residue while the same symbol without hyphens denotes the free amino acid.
Are Gly-Ala and Ala-Gly the same peptide?
No. Both contain one glycine residue, one alanine residue, and one peptide bond, but the connectivity differs. In Gly-Ala the glycine residue is N-terminal and its carbonyl carbon is bonded to the alanine nitrogen. In Ala-Gly the alanine residue is N-terminal and the bond runs the other way. The nomenclature names them glycylalanine and alanylglycine, which are different compounds.
Is there one exact amino-acid count that separates a peptide from a protein?
No. Prefixes such as di-, tri-, and tetra- state exact residue counts, but oligopeptide, polypeptide, and protein are approximate size conventions. The IUPAC-IUB recommendations describe those length terms as conventions and note that authors disagree about the boundaries. A document that uses the word peptide or protein has told you roughly how it thinks about size, not an exact residue count.
Does an N-terminal modification mean a peptide has no N-terminus?
No. The N-terminal residue is the residue whose amino nitrogen is not part of a peptide bond, and it keeps that position whether or not the nitrogen carries a modifying group. The recommendations retain the terms N-terminal and C-terminal for modified ends. A modification should appear as an annotation on the terminal residue, and its absence from a written sequence does not prove the terminus is unmodified.
Does a one-letter sequence tell me every chemical detail?
No. A one-letter sequence gives residue identity and order in the most compact form and nothing else. It does not encode D or L configuration, R or S descriptors, terminal state, non-alpha linkages, or residue modifications. Placeholders such as B, Z, and X mark unresolved or unknown residues. A complete description adds explicit annotations for each of those items, using a standard such as ProForma where modifications are present.
Sources and scope notes
The definitions, residue and naming conventions, stereochemical descriptors, symbol rules, and one-letter conventions on this page are paraphrased from the IUPAC-IUB Joint Commission on Biochemical Nomenclature recommendations on amino acids and peptides, as hosted by Queen Mary University of London, and from the official ProForma project page maintained by the HUPO Proteomics Standards Initiative. Each was read on September 20, 2026, Pacific time. The recommendations are foundational nomenclature documents dating from 1983 with later addenda; the reading date is not their publication date.
Representation limits: the sketches on this page are teaching diagrams drawn in the neutral form used by nomenclature tables. They omit stereochemistry, some hydrogens, and any ionization state, and they carry no analytical measurement, batch record, or activity information. No statement here infers use suitability, biological activity, or sample identity from a name, formula, or sequence.
Research-only scope: this page describes molecular description and notation for laboratory research documentation. It is not a synthesis procedure, a purchasing comparison, a certificate interpretation guide, or a description of any specific research material or lot.
- IUPAC-IUB JCBN, Nomenclature and Symbolism for Amino Acids and Peptides (Recommendations 1983), sections 3AA-11 to 3AA-13: peptides, residues, and naming Hosted by Queen Mary University of London. Read September 20, 2026.
- Same recommendations, Table 1 and section 3AA-2.2.5: amino acid structural formulas, the main chain, and side chains Read September 20, 2026.
- Same recommendations, sections 3AA-3 to 3AA-5: configuration descriptors, optical rotation, and racemates Read September 20, 2026.
- Same recommendations, sections 3AA-14 to 3AA-16: three-letter symbols and residue notation Read September 20, 2026.
- Same recommendations, sections 3AA-20 and 3AA-21: one-letter symbols and sequence representation Read September 20, 2026, with linked later addenda.
- HUPO Proteomics Standards Initiative, ProForma: version table listing 2.0 (final February 3, 2022) and 2.1 (final June 9, 2026) Landing page read September 20, 2026. The full specification was not used.
